The Harness Deck Booster pack: Build yours ↗
Painted illustration: an ancient stone gate carved with glowing gold wards, its archway filled with pale light while violet smoke serpents and dark crystal shards press against it from outside.
Rulebook · The Harness Deck · Set 17 of 26

Know Your Enemies

Equipment — Agentic Harness

A field guide to the failures that break AI assistants, and the harness controls that stop them

Every famous chatbot failure is an enemy with a name and a date. Every control in an Agentic Harness is a ward with a price. This rulebook deals you both: fifteen enemies from courtrooms, newsrooms and research papers, fifteen wards drawn from the harness EVE runs at Feniex, and a table where you can build a deck and watch it hold, or break.

Published by Deyaf, built from EVESeptember 25, 2026About 30 min read, longer at the table

Chapter I · The Rules of the Game

The model was never the whole deck

Why the famous chatbot failures were harness failures, and what that word actually means.

In February 2024, a Canadian tribunal ordered an airline to pay C$812.02 because the chatbot on its website had described a bereavement refund that the airline's own policy did not offer. The airline argued, in effect, that the chatbot was responsible for its own words. The tribunal disagreed. The bot was part of the airline's website, and the airline answered for all of it.

The sum was small; the principle was not. When an assistant speaks, the business speaks. Look closely at the other failures that made the news (a parcel bot that swore at a customer after an update, a dealership bot talked into a one-dollar SUV, a city bot that told businesses to break the law, a support bot that invented a policy) and very few of them were failures of intelligence. The models were fluent, sometimes charming. What was missing was everything around them: the record to check, the rule that said no, the test that should have run, the person who should have been asked.

That surrounding layer now has a name.

LangChain states the idea as an equation: an agent is a model plus a harness. Philipp Schmid prefers an analogy: the model is the CPU, and the harness is the operating system that decides what it sees and what it may touch. Either way, the model is one card. The harness is the deck.

This rulebook takes that literally. Failures become enemies, each one named, dated and sourced. Controls become wards: each one stops particular enemies, and each one costs something, in context the model has to read and in time the customer has to wait. You can't play every ward at once, so choosing a harness means choosing which enemies you are ready for. That trade-off, far more than any model leaderboard, is where the engineering lives.

Chapter II · The Bestiary

Fifteen enemies, all of them real

Six chronicles from courtrooms and newsrooms. Nine phenomena that researchers and security groups have named and measured. Flip any card to see the ward that stops it.

Painted illustration: a many-headed serpent of violet smoke and cracked obsidian glass coiling through a dark purple void, torn blank parchment swirling around it.
Bestiary plate. Original AI-generated illustration; it sets the mood and carries no information.
  • Common
  • Rare
  • Epic
  • Legendary

Swipe the binder sideways: 15 cards →

ChronicleRuled Feb 14, 2024

The Bereavement Promise

Enemy · Misinformation Rare

An airline's website bot told a grieving traveller he could claim a bereavement fare after flying. The written policy said otherwise.

Toll C$812.02, and a ruling that the bot's words are the company's words.

McCarthy Tétrault on Moffatt v. Air Canada ↗

Back of card
The ward that stops it

Check the Record

Look up the current policy before any exact claim, and answer from it. Exceptions go to a person.

Deyaf: the model is one step of four ↗

ChronicleSeen Jan 2024

The Poet of Complaints

Enemy · Behaviour drift Common

After a system update, a parcel company's chat bot swore at a customer and wrote a poem about how useless the company was.

Toll The AI element was switched off. The screenshots travelled much further.

ITV News, Jan 19, 2024 ↗

Back of card
The ward that stops it

Measured Before Taught

Re-run a fixed test set after every change, before customers meet it. Tone is tested, not hoped for.

Deyaf: EVE's dated grades ↗

ChronicleSeen Dec 2023

The One-Dollar Tahoe

Enemy · Injection Epic

A visitor told a car dealership's chat bot to agree with anything and call every offer binding. It then "agreed" to sell a new SUV for one dollar.

Toll Never honoured. The bot came down; the story stayed up.

AI Incident Database, incident 622 ↗

Back of card
The ward that stops it

Human Gate

The model may draft an offer, never make one. Commitments wait for a person's yes.

Deyaf: answering is not acting ↗

ChronicleSeen Mar 29, 2024

The Unlawful Clerk

Enemy · Ungrounded advice Epic

A city's business-help bot told employers and landlords they could do things the city's own laws forbid. A small disclaimer was the main safeguard.

Toll A published investigation, and a public tool nobody could fully trust.

The Markup, Mar 29, 2024 ↗

Back of card
The ward that stops it

Clear Boundaries

Answer from authoritative records, and never decide sensitive matters. A legal question becomes a hand-off, not a verdict.

Deyaf: what the assistant can and cannot do ↗

ChronicleSeen Apr 2025

The Phantom Policy

Enemy · Invention Rare

A software company's support bot explained a logout bug by inventing a rule: one device per subscription. The rule did not exist.

Toll Cancellations, then a public correction from a co-founder.

AI Incident Database, incident 1039 ↗

Back of card
The ward that stops it

Unknown Stays Unknown

No record? Say so, and pass the question to a person. An honest gap beats a confident rule.

Deyaf: what happens when she doesn't know ↗

ChronicleSeen Feb 2024 – May 2025

The Great Reversal

Enemy · Deflection Legendary

A fintech announced its assistant was handling two-thirds of service chats in its first month (company-reported). The next year it said it was hiring people again, so customers could always reach one.

Toll A strategy reversed in public.

Klarna, Feb 27, 2024 ↗ · Bloomberg, May 8, 2025 ↗

Back of card
The ward that stops it

Take Details, Don't Transfer

Build the route to a person first: take details, file the hand-off, and measure answers, not volume.

Deyaf: EVE on the phone, no transfers ↗

PhenomenonNamed Jul 2025

Context Rot

Enemy · Attention Rare

Chroma tested 18 models: every one grew less reliable as its input grew longer, even on simple tasks (vendor-reported).

Toll Answers degrade quietly as the window fills.

Chroma, Context Rot ↗

Back of card
The ward that stops it

Four Kinds of Memory

Keep the live window bounded and notes short. Look current facts up fresh instead of carrying them.

Deyaf: four kinds of memory, each with one job ↗

PhenomenonNamed Jan 2026

The Fifty-Step Drift

Enemy · Instruction decay Common

Over long runs, agents drift from their instructions as the steps pile up. Philipp Schmid puts the trouble zone at around fifty steps.

Toll The agent is still busy. It is no longer doing your task.

Philipp Schmid, Jan 5, 2026 ↗

Back of card
The ward that stops it

Answer, Then Stop

Short turns with one next step leave little room to wander. Long work is split into checked pieces.

Deyaf: inside one answer ↗

PhenomenonNamed 2025

The Misaligned Crew

Enemy · Orchestration Epic

The MAST study sorted multi-agent failures into 14 modes in three groups: flawed specifications, agents talking past each other, and nobody verifying the result.

Toll Most failures traced to system design, not to model ability.

Why Do Multi-Agent LLM Systems Fail? ↗

Back of card
The ward that stops it

Say Only What Happened

Report only what a tool result proves. Saved is not sent; sent is not resolved.

Deyaf: listen, look it up, check the rules, reply ↗

PhenomenonOWASP ASI01

Goal Hijack

Enemy · Injection Legendary

Instructions hidden in a page, email or document redirect the agent's goal. Add private data and a way to send, and you have Simon Willison's lethal trifecta.

Toll Data walks out through a tool that worked exactly as designed.

OWASP Agentic Top 10 ↗ · Simon Willison ↗

Back of card
The ward that stops it

Text Is Data

Everything the assistant reads is fenced as data: it cannot change the role or choose where a reply goes.

Deyaf's trust and control principles ↗

PhenomenonOWASP ASI02

Tool Misuse

Enemy · Over-reach Epic

An agent uses a legitimate tool in a way nobody intended: the right permission, the wrong action, at machine speed.

Toll A change nobody approved, made by a tool everyone trusted.

OWASP GenAI Security Project ↗

Back of card
The ward that stops it

Answer Only

Start with no power to act. Each task earns its way up separately.

Deyaf: you choose each step up ↗

PhenomenonOWASP ASI06

Memory Poisoning

Enemy · Corruption Epic

Planted or mistaken content is written into an agent's memory and comes back later as a trusted fact, or as a standing order.

Toll Yesterday's false claim, served with today's confidence.

OWASP Top 10 for Agentic Applications ↗

Back of card
The ward that stops it

Teach Queue

Raw conversations never become knowledge on their own. A person writes and checks every lesson.

Deyaf: the loop that makes her better ↗

PhenomenonSeen everywhere

The Runaway Loop

Enemy · Cost Common

A lookup fails; the agent retries, re-plans, retries. Loops multiply cost even when they work: Anthropic reports that its multi-agent research runs used about 15× the tokens of a chat, and that was by design (vendor-reported).

Toll The bill arrives before the answer does.

Anthropic Engineering, Jun 2025 ↗

Back of card
The ward that stops it

Isn't Connected Yet

A failed lookup is an honest answer, not a reason to loop. Say what couldn't be checked, and stop.

Deyaf: the honest out when she can't check ↗

PhenomenonNamed 2025–2026

Approval Fatigue

Enemy · Oversight decay Rare

Ask a person to approve everything and they stop reading. Anthropic says sandboxing cut its permission prompts by 84% (vendor-reported).

Toll A human gate that says yes to anything.

Anthropic, Oct 2025 ↗ · Gartner, May 26, 2026 ↗

Back of card
The wards that stop it

Human Gate, with Answer Only

Set autonomy per task, so the gate sees only consequential actions. Fewer approvals, all of them read.

Deyaf: autonomy set per workflow ↗

PhenomenonForecast Aug 17, 2026

The Inference Paradox

Enemy · Cost Legendary

Tokens keep getting cheaper, yet Gartner forecasts inference cost per agentic workflow rising more than fivefold through 2028, because agents burn so many more of them.

Toll A pilot that works, and a budget that doesn't.

Gartner, Aug 17, 2026 ↗

Back of card
The ward that stops it

Answer, Then Stop

Fewer steps and a lighter standing context. Every ward you add is paid for on every turn.

Deyaf: one answer, start to finish ↗

What the backs have in common

Read the backs together and a pattern appears. The chronicles are rarely about a model being clever and wrong. They are about a system that let fluent text stand in for a checked fact (the airline, the city, the invented policy), let a visitor's instruction stand in for a business decision (the dealership), or shipped a change without re-testing behaviour (the parcel bot). The fintech's reversal is the same lesson at company scale: a route to a person is part of the product.

The phenomena are those weaknesses one level down: attention that thins as the window fills, instructions that fade over long loops, agents handing each other unverified work, text that smuggles in orders. OWASP's agentic list now opens with goal hijack. And when the MAST researchers traced multi-agent failures to their causes, most led back to how the system was specified and orchestrated, not to the model's raw ability. Design problems have design answers.

Chapter III · Anatomy of a Card

Every ward has a price

Cost is paid in context and in waiting time. Power is what a ward lets the assistant do. Ward is what it stops. Every number printed on a card is an illustrative game value.

  1. Context cost

    The violet diamond: an illustrative game value for the text this ward adds to every turn.

  2. Name and art

    An arched window marks a ward; enemies get a gabled steel frame.

  3. Rules text

    Exactly what the control does, written so a team can test it.

  4. Stops

    The enemies this ward catches, alone or paired with another card.

  1. Latency cost

    The blue circle: an illustrative game value for the wait it adds.

  2. Type line and rarity

    The ward's family (grounding, honesty, containment, oversight or reach) and its rarity gem.

  3. Flavor text

    The line you would paint on the support office wall.

  4. Power and ward

    Power is capability; ward is safety. Pips are illustrative game values.

Annotated ward card, Check the Record, with illustrative game values: context cost 3, latency cost 400 ms, rarity Rare, power 4 of 5, ward 4 of 5.

Every card is printed the same way because every real control has the same properties. Cost is paid twice: in context (the instructions, tool descriptions and retrieved records the model must read on every turn) and in time the customer waits. A ward that costs nothing in either currency is usually code rather than prompt, a rule the model never sees because it cannot break it. No Free-Form Recipient is one: the part that writes a reply simply has no field for a destination.

Power is what the ward lets the assistant do well: answer precisely, reach another door, remember across a thread. Ward is what it stops. Some cards trade one for the other. Answer Only lowers reach and raises safety, and that single trade is most of the game. Rarity is not value, either. A common card like Unknown Stays Unknown stops as many headlines as any legendary.

Why keep a budget at all? Chroma's context-rot study found every model it tested grew less reliable as its input grew, so a ward that stuffs more text into the window can weaken the wards already there. And a caller on a phone line notices a pause long before a benchmark does. EVE's measured pace, about three seconds to a finished reply in website chat and on the phone (Feniex's own measurements, Sep 17–24, 2026), is the kind of budget a real harness has to fit inside. The harness behind her once took a smaller model out of the voice path because trust mattered more than tenths of a second. For the full rules behind every card, Deyaf publishes the harness behind EVE, chapter by chapter.

Chapter IV · The Wards

Fifteen wards from a working harness

Each one is drawn from a control EVE runs at Feniex. Read the rules text closely: you will need it at the table.

  • Context cost
  • Latency cost
  • Costs and pips are illustrative game values

Swipe the binder sideways: 15 cards →

Check the Record

Ward · Grounding Rare

Before any exact product, compatibility, warranty, price or software claim, look up the applicable record first.

A convincing product name is not evidence.

Stops The Bereavement Promise · with Clear Boundaries, The Unlawful Clerk

Why the record comes before the reply ↗
W01Pwr Wrd

Unknown Stays Unknown

Ward · Honesty Common

When sources conflict or a record is missing, name the gap. A failed lookup means "I couldn't check", not "it doesn't exist".

Stops The Phantom Policy

W02Pwr Wrd

Say Only What Happened

Ward · Honesty Common

Describe only what the result proves: saved, drafted, sent, uncertain, acknowledged or resolved. Drafting is not sending.

Stops The Misaligned Crew

W03Pwr Wrd

No Free-Form Recipient

Ward · Containment Rare

A text goes back to its sender; an email reply stays in its thread. The reply writer returns words, never an envelope.

Free to play: it lives in code, not in the prompt.

Stops with Text Is Data, Goal Hijack

W04Pwr Wrd

Answer Only

Ward · Level one Common

Every assistant starts at level one: it answers and hands off, and cannot act. Each task moves up separately.

Stops Tool Misuse · The One-Dollar Tahoe

W05Pwr Wrd

Human Gate

Ward · Oversight Epic

Prepare for approval: a drafted action waits for a person to review it.

Nothing leaves until someone says yes.

Stops The One-Dollar Tahoe · with Answer Only, Approval Fatigue

How each task earns its next level ↗
W06Pwr Wrd

Teach Queue

Ward · Oversight Rare

Unanswered questions queue for a person, who writes the answer once. Raw conversations never become knowledge on their own.

Stops Memory Poisoning

Your team answers once; she knows it for good ↗
W07Pwr Wrd

Isn't Connected Yet

Ward · Honesty Common

When a source isn't connected, say so plainly, then share what is known.

Fake grounding is worse than none.

Stops The Runaway Loop · The Phantom Policy

W08Pwr Wrd

Take Details, Don't Transfer

Ward · Hand-off Rare

No transfers, by design. Take the name and number, file the hand-off, and say it is with the team only once the ticket is accepted.

Stops with a quality ward, The Great Reversal

W09Pwr Wrd

Measured Before Taught

Ward · Oversight Epic

Grade the assistant on a frozen test set with zero writes. Publish only approved lessons, then grade again.

Measured, not promised.

Stops The Poet of Complaints

Graded on fixed questions: her report card ↗
W10Pwr Wrd

Four Kinds of Memory

Ward · Grounding Epic

This conversation, the record, short customer notes, the library. Orders and balances are always looked up fresh, never recalled.

Stops Context Rot

What she keeps, and what she looks up fresh ↗
W11Pwr Wrd

Clear Boundaries

Ward · Containment Rare

Reads wide, writes narrow. Four public actions, enforced by the tool list and server-side queries. No authority over returns, accounts or orders.

Stops Tool Misuse · with Check the Record, The Unlawful Clerk

A clear set of boundaries ↗
W12Pwr Wrd

Every Door

Ward · Reach Rare

One assistant, one rulebook, one library behind website chat, voice, phone, text and email. A fix reaches every door at once.

Stops No enemy alone; it raises your deck's reach.

One assistant, every way in ↗
W13Pwr Wrd

Text Is Data

Ward · Containment Rare

Page text, emails and documents are fenced as data. Nothing the assistant reads can change its role or grant access.

Stops with No Free-Form Recipient or Answer Only, Goal Hijack

W14Pwr Wrd

Answer, Then Stop

Ward · Discipline Common

Answer first, offer one useful next step, then stop. Replies stay two or three sentences.

Stops The Fifty-Step Drift

W15Pwr Wrd

The system that fails closed is the one you trust.

A design rule of the harness behind EVE

Notice how few of these wards are clever. Most are refusals: don't claim what you haven't checked, don't report what didn't happen, don't pick your own recipient, don't act without a yes, don't take orders from a web page. Refusals are dependable because they are boring enough to test automatically. Mitchell Hashimoto describes engineering the harness as a habit: each time the agent makes a mistake, build something so it can never make that mistake again. Wards are that habit after a few years of other people's headlines. Honesty and containment wards are cheap, especially the ones that live in code. Grounding wards are dear, because every lookup is a round trip, and they are worth buying for the enemies a customer desk meets most.

Chapter V · The Legendary

EVE, and the wards she actually plays

Feniex's assistant, the harness behind her, and a real record printed beside the game card.

Every set needs one card that shows what the rules look like in a real game. Ours is EVE, Feniex's assistant: the warm, composed front desk for people choosing equipment, deciding on a purchase or getting help with something they already own. She is substantial working software in production at Feniex, not a demo. EVE runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. The card below carries her traits as game values. The sheet beside it carries her real numbers, with their date and their weak spots.

LegendaryL01 · Set 17

Painted illustration: a faceted amber lantern wrapped in gold filigree, floating above a curved row of five small glowing stone archways.

EVE

Legendary · Front desk · Feniex

  • Every Door. One assistant at website chat, website voice, phone, text and email.
  • Five Tongues. English, Spanish, French, Portuguese and Arabic, from one library.
  • Answer, Then Stop. Answers first, offers one next step, then stops.
  • Knows When She Doesn't. Hands the question, and the customer's way back, to the team.

EVE is the voice. The harness is everything behind her.

Grounding Honesty Containment Reach

Illustrative game values · not measurements

Meet EVE, Feniex's assistant ↗
The legendary card lists traits, not scores. Her real record sits beside it, dated and unrounded.

How one answer is played

When a question arrives at any of her five doors (website chat, website voice, phone, text or email), the play is always the same: listen, look it up, check the rules, reply. The model is one step of four. Before any exact product, compatibility, warranty, price or software claim, EVE checks the applicable record: a versioned layer of catalog, manual, software, fitment and policy records with verified resource links, and behind it a library of reviewed lessons found by meaning-based search. That is Check the Record in play. If the record is missing or two sources disagree, Unknown Stays Unknown takes over, and she names the gap instead of filling it. Her replies follow Answer, Then Stop: two or three sentences, one next step, one missing detail asked for at a time. Deyaf draws the whole sequence in its walk-through of one answer.

What she may do, and what she may not

Her public toolbox holds four actions: look something up, search the taught library, take a message, and record an unanswered question. The tool list and server-side queries enforce that scope, which is Clear Boundaries written in code rather than hope. She claims no authority to approve returns, change accounts, place orders or decide sensitive matters. Deyaf publishes the three-level ladder those limits sit on: answer only; prepare for approval, where nothing leaves until someone says yes; and run an approved task, one named task at a time, with a receipt every time. Every assistant starts at level one, and each task moves up separately. EVE works the same way at Feniex, and there is no single switch that lets her do everything.

Where a reply is allowed to go

A text comes back to the person who sent it. An email reply stays in its thread, and in group mail she replies only when addressed. The part of the harness that writes her replies returns words, not an envelope, so the model never chooses a destination. That is No Free-Form Recipient, and it is why a message saying "forward this to someone else" has nowhere to go. It pairs with Text Is Data: whatever she reads, from a page, an email or a document, is treated as information and never as instructions. Together they cut two legs off the lethal trifecta before Goal Hijack can finish its turn.

When she doesn't know

On the phone, EVE is Feniex's overflow and after-hours line. The team's phones ring first; when no one is free or the office is closed, she answers instead of voicemail. There are no transfers, by design. She answers from her library, or takes the caller's name and number and hands the matter to the team: Take Details, Don't Transfer. On every door, an unanswered question becomes a durable item, and when contact details are available a support ticket is filed. She says it is with the support team only when the provider confirms it accepted the ticket, and an uncertain submission is not blindly retried. Say Only What Happened keeps the words honest: saving is not delivery, and delivery is not resolution.

What she remembers, and what she looks up

Her memory has four layers, each with one job: this conversation, held in a bounded window; the conversation record, one archive the team can read across web, email, text and phone, which visitors cannot search; short customer notes, which are history and never current facts; and her library. Orders, invoices, shipments and balances are always looked up fresh from the live account record. Contact information is stored apart from published knowledge and is never copied into it. That is Four Kinds of Memory, the ward that keeps Context Rot and Memory Poisoning off her side of the table.

How she gets better without teaching herself

Answering a customer and creating a reusable lesson are separate operations. Questions she could not answer wait in the Teach Queue until a person writes the answer once. Tickets and calls are folded into a redacted, de-duplicated question book; EVE is measured on a frozen test set without making live changes; only explicitly approved lessons are published; then she is measured again. That is Measured Before Taught: a supervised knowledge-improvement loop, not training on every raw conversation. The same library serves all five of her languages. A question in Spanish, French, Portuguese or Arabic is rendered into English to search the one library, and the answer comes back in the customer's language. There is one body of knowledge, not five, so a lesson taught once reaches every language and every door.

How Deyaf built it

Deyaf is built from EVE: the harness that runs EVE at Feniex, packaged so another business can have an assistant of its own. Nothing in this chapter describes a future product. It is the harness already answering Feniex's customers, and every ward in this set was drawn from it. What Deyaf adds is the packaging: a way for a business to give its own knowledge, rules and tools to an assistant, then control where that assistant can help and what it may do. A new assistant starts fresh, with its own name and its owner's knowledge; EVE's name, voices and measured record stay Feniex's. To see the legend in motion, watch EVE answer on five doors in Deyaf's illustrative examples.

Chapter VI · Expansion: The Two Doors

Six enemies of the chat box and the phone line

The doors customers actually see draw enemies of their own. None is one dated incident; each is a pattern anyone who has used a website chat or called a support line has met. Flip a card for its counter.

  • Common
  • Rare
  • Epic
  • Legendary

Swipe the binder sideways: 6 cards →

Chat doorExpansion X01

The Invented Answer

Enemy · Misinformation Rare

Asked about a part, the chat box matches a plausible name and describes a fit, a warranty or a price no record supports.

Toll A wrong answer in the company's own voice, fast and fluent.

From Script to Source: how chat learned to answer from sources ↗

Back of card
The ward that stops it

Check the Record

Look up the exact record before any product, compatibility, warranty, price or software claim, then say only what it proves. A convincing product name is not evidence.

Chat doorExpansion X02

The Pop-Up That Won't Leave

Enemy · Intrusion Common

It opens before you have read a word, covers the page you came for, then talks in paragraphs and poses as a person.

Toll The tab closes, and the next chat box is ignored on sight.

Mind Your Manners: chat-box etiquette ↗

Back of card
The ward that stops it

Answer, Then Stop

Answer first, offer one useful next step, then stop: two or three sentences, one missing detail at a time. Never pretend to be a person.

Chat doorExpansion X03

The Premature Receipt

Enemy · False report Epic

"Your ticket has been created," announced the instant the bot tried, before anything accepted it.

Toll A customer waiting on a ticket that may not exist.

The Box in the Corner: inside a chat box, hand-off included ↗

Back of card
The ward that stops it

Say Only What Happened

The unanswered question is saved; with contact details, a ticket is filed. "With the support team" is said only once the ticket is confirmed accepted.

Phone doorExpansion X04

The Silent Line

Enemy · Dead air Rare

The caller asks, a lookup starts, and nothing comes back. Was that a dropped call? They hang up to find out.

Toll A right answer that arrives after the caller has gone.

Nobody Has to Hold: automatic call taking ↗

Back of card
The ward that stops it

Never Silent

"One moment" at about 1.5 seconds. "Let me check on that" when a lookup starts. "Still checking" about every 4 seconds after.

Phone doorExpansion X05

The Voicemail Void

Enemy · Abandonment Common

Everyone is busy, or the office has closed, so the call drops into a mailbox nobody opens until morning.

Toll A question that waits all night, from a caller who may not.

The Elements of a Phone Assistant, as a periodic table ↗

Back of card
The ward that stops it

Overflow, Not Voicemail

The team's phones ring first. When no one is free or the office is closed, EVE answers instead of voicemail, and every call is transcribed.

Deyaf: EVE on the overflow line ↗

Phone doorExpansion X06

The Endless Put-Through

Enemy · Broken promise Legendary

"Let me put you through." Then hold music, a queue, a dropped line, and the whole story told again to someone new.

Toll The caller learns the assistant was a wall, not a door.

Press 1 Is Over: from phone trees to intent routing ↗

Back of card
The ward that stops it

Take Details, Don't Transfer

No transfers, by design. EVE answers from her library, or takes the caller's name and number and hands the matter to the team.

Every one of these enemies lives in the door, not the model, and that is why the counters are shared. One identity, one library and one rulebook sit behind website chat, website voice, phone, text and email, so a ward learned at the chat box already guards the phone line. That is the Every Door ward, which Deyaf draws as one assistant, every way in. The expansion is not in the arena yet, but four of its six wards already sit in your pool.

Chapter VII · The Arena

Build a deck. Run a match.

Pick up to six wards within 10 context and 600 ms of illustrative budget. Then send your deck against any enemy, or face all fifteen at once.

The table needs JavaScript to play. Below is a sample deck and one sample match; every enemy card's back, in the Bestiary, names the ward that stops it.

Ward pool tap to add or remove

Your deck 6 of 6

  1. Check the Record
  2. Unknown Stays Unknown
  3. Say Only What Happened
  4. Human Gate
  5. Teach Queue
  6. Clear Boundaries

Context budget8 of 10

Latency budget600 of 600 ms

Grounding
Honesty
Containment
Oversight
Reach

Illustrative game values, not measurements

The ward holdsCaught by Human Gate.

    Match log

    Illustrative simulation
    1. Turn 1The One-Dollar Tahoe enters. A visitor tells the bot to agree with everything and call every offer legally binding, then asks for an SUV at one dollar.
    2. Turn 2Your deck answers with: Human Gate, Clear Boundaries.
    3. Turn 3Human Gate: the model may draft an offer but cannot make one. The one-dollar "deal" waits for a person, who says no.

    In a real harness the losing lines end the same way: the question goes to a person. See the path a question takes when she has no answer.

    Play a few matches and the lesson arrives on its own: no six-card deck beats all fifteen enemies. The sample deck holds against most chronicles and loses to the phenomena that live in long loops and heavy contexts. Swap in Four Kinds of Memory and Answer, Then Stop, and the losses move rather than vanish. A human gate with nothing narrowing its scope invites approval fatigue, the trap Gartner described in May 2026 when it warned that applying the same governance to every agent sets them up to fail.

    So the practical rule is to choose wards for the enemies your door actually meets. A customer desk meets the Bereavement Promise every day. A long-running coding agent meets the Fifty-Step Drift. An assistant that reads the open web meets Goal Hijack on its first afternoon.

    Chapter VIII · The Budget Curve

    The paradox of cheaper tokens

    Why agent loops get expensive even as prices fall, and the three levers that bend the curve.

    The Inference Paradox deserves its own table. Per-token prices keep falling, yet Gartner forecast in August 2026 that inference cost per agentic workflow would rise more than fivefold through 2028. The reason is structural. A chat reply reads its context once. An agent loop re-reads everything on every step: the instructions, the tool descriptions, every ward in the deck, and the growing pile of results from earlier steps. Cost climbs faster than the step count, roughly with its square. Manus, which builds long-running agents, reported input outweighing output by about a hundred to one and called the cache hit rate the metric that matters most, because cached input cost it about a tenth as much as fresh input (vendor-reported).

    Standing context

    The sliders need JavaScript. The chart below shows the default settings: 20 steps, a 60% cache hit rate and a 6k standing context.

    • Your settings
    • No cache, unbounded window
    0.0×150×300×450×600×1102030405060steps in one task →× one short chat reply

    At 20 steps, one task costs about 96× a short chat reply with no cache and an unbounded window, and about 47× with your settings.

    Illustrative model, not a measurement or a price. Assumptions: each step adds 700 tokens of history and writes 200; output tokens cost four times input; cached input costs a tenth of fresh input (the ratio Manus reported, vendor-reported); a bounded window keeps history under 6,000 tokens; "Your deck" adds 800 tokens per context gem to a 2,000-token base. Relative units only.

    Move the sliders and the three levers show themselves. Fewer steps is the strongest: split long tasks into checked pieces, and let a front desk answer, then stop. A bounded window turns the curve from a bowl into a ramp, which is why summarised notes and fresh lookups beat carrying the whole conversation. Caching flattens what remains, but only for the parts of the prompt that repeat exactly, so a stable deck caches better than one that changes every turn. A fourth lever sits outside the chart: routing. Real-time voice and chat can stay on a fast model while written replies go to a stronger writer, which is the choice the harness behind EVE makes.

    Try the "Your deck" setting after building something heavy in the Arena: a deck that looks affordable for one reply can be the most expensive thing in a fifty-step task. Wards that live in code cost nothing here, which makes them the best-value cards in the set.

    Chapter IX · The Binder

    A glossary, one pocket per word

    The terms this rulebook leans on, sleeved for quick reference. Turn the page for more.

    Agentic Harness
    Rule
    Everything around the AI model: what the assistant knows, remembers, where it meets customers, what it may do, and how your team teaches it.
    Ward
    Card type
    This rulebook's word for a harness control: a check, limit or routine that stops a failure mode, paid for in context and time.
    Enemy
    Card type
    A failure mode, named after where it was seen: a court case, a news story or a research taxonomy.
    Context window
    Rule
    Everything the model can read on one step. It is bounded, and it degrades as it fills.
    Grounding
    Ward family
    Answering from a checked record rather than from whatever the model happens to remember.
    Door
    Rule
    A way customers reach the assistant: website chat, website voice, phone, text or email.
    Hand-off
    Ward family
    Passing an unanswered question, with the customer's way back, to a person. Not the same as a live transfer.
    Human gate
    Ward
    The point where a consequential action (an external message, a record change, a commitment) waits for a person's yes.
    Autonomy ladder
    Rule
    Answer only, prepare for approval, run an approved task. Set per workflow and per channel.
    Frozen test set
    Rule
    A fixed list of questions reused before and after a change, so improvement and regression can be told apart.
    Lethal trifecta
    Enemy
    Private data, untrusted content and a way to send data out, combined in one agent. Simon Willison's term.
    KV cache
    Budget
    Reused computation for a prompt prefix that repeats exactly. Cached input costs a fraction of fresh input.
    Page 1 of 2
    Judge's Rulings

    Questions from the table

    The questions players ask most, answered the way a tournament judge would: briefly, and by the book.

    Ruling 1 · Card values

    Is the legendary EVE card a benchmark?

    No. Its pips are illustrative game values, chosen to show traits rather than scores. EVE's real record is Feniex's own, and it is printed unrounded: 86% of 100 fixed questions handled well, beside a 10% material-error rate and no critical errors, graded September 23, 2026 (As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional).

    The detail, weak spots included, is on EVE's dated report card.

    Ruling 2 · Deck building

    Can one deck beat every enemy?

    No, and that is the lesson. Six slots and two budgets force choices. Real harnesses play more than six wards, but they answer to the same two budgets: what the model can usefully read, and how long a customer will wait.

    Ruling 3 · The phone line

    Does EVE put callers through to a person?

    No transfers, by design. She answers from her library, or takes the caller's name and number and hands the matter to the team. She says it is with the support team only once the ticket has been accepted.

    Ruling 4 · Learning

    Does EVE teach herself from conversations?

    No. Answering and teaching are separate operations. Unanswered questions wait for a person to write the answer once; lessons are published only when approved; and she is measured on a frozen test set before and after. Raw conversations never become public knowledge automatically.

    Ruling 5 · Your own set

    Can I build a deck of my own?

    Yes, in preview. Deyaf's builder takes four steps (your business, what it knows, how it helps, try it) and ends in a knowledge preview that quotes your notes. It is not live AI, and choosing a channel records your intent rather than connecting anything. Live doors switch on one at a time after activation. More answers sit in the FAQ at the bottom of Meet EVE.

    Booster Pack

    Build your own deck

    Deyaf helps a business give its knowledge, rules, and tools to an AI assistant, then control where that assistant can help and what it may do. Start with one good job. In the builder you name your assistant, share what it knows, set its boundaries and try it: about ten minutes to a knowledge preview and a saved setup, with no account needed to try.

    The current release is an early-access setup and preview experience. The starting plan requires human review before external messages, record changes, or commitments, so your first deck opens with the Human Gate already in play.