Answers, Then Stops Site 14 · an essay on restraint Build yours ↗

Answers, then stops.

An essay on restraint in the

Agentic Harness

A good reply from an AI assistant is often three sentences long: the answer, where it came from, and one next step. Then nothing. This essay is about the engineering behind that nothing, and about Eve, the assistant who works that way at Feniex.

Deyaf by Feniex · 25 September 2026 · About 14 minutes to read slowly

A single ink brushstroke sweeps across pale paper, thinning to a dry trail and stopping, with one small vermilion dot beyond its end.

The shape of an answer

Ask a well-run front desk a question and notice what does not happen. Nobody recites the catalog. Nobody guesses a date to make you feel better. You get the answer, the page it came from and one thing to do next, and then the person stops talking and lets you decide. The pause at the end is not rudeness. It is the part of the service that leaves room for you.

AI assistants find that pause hard. A language model is built to continue: give it a question and it produces the most plausible next word, and the next, for as long as it is allowed. Fluency is its native gift, and fluency does not know when it has run out of evidence. Left alone, a model fills silence with hedges, a second offer, a restated question, a cheerful guess.

So the stopping has to come from somewhere else. It comes from the system around the model, which the industry has started calling a harness. This library uses Deyaf's definition:

An Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.

Deyaf's definition

Restraint lives in that "everything". It is written into the rules the model reads, the records it is allowed to quote, the verbs it is allowed to use and the people who review what it learns. None of it is visible in a single reply, and all of it decides where that reply ends.

Eve is a working example. She is Feniex's assistant, the warm and composed public face that helps people choose suitable equipment, make a buying decision or get support for what they already own. Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. Her manner is a specification, not a mood. She answers first, gives one useful next step, then stops; her replies are usually two or three sentences. She asks for one missing detail at a time and accepts corrections. She respects a customer's budget and recommends the simpler option when it is enough. She acknowledges frustration once, and she never tacks a sales pitch onto a support answer.

None of that is decoration. Anthropic's engineers, writing about the agents that hold up in production, recommend the simplest solution that does the job, adding complexity only when it demonstrably helps. The same discipline applies inside a single reply. Every extra sentence is a new claim the customer has to trust, and every claim needs somewhere to come from.

Illustrative · not measured

Answers, then stops three sentences and a stop

  1. the answer
  2. where it came from
  3. one next step
the space after the answer, left for the customer

Keeps going the same answer, buried

  1. a warm-up
  2. the question, restated
  3. the answer
  4. a hedge
  5. a guess at a date
  6. a second offer
  7. a pitch
  8. three options
  9. another hedge
  10. "anything else?"
Two replies to one question, drawn as shapes. Each bar is a sentence. The reply on the left carries the same answer as the one on the right, with nothing the customer has to verify or wade through, and it ends.

Deyaf ↗Meet Eve, and how she talks

Check the record

The first rule of Eve's evidence contract is short enough to fit on a card. Before any exact product, compatibility, warranty, price or software claim, she checks the applicable record. Her operating rules give the reason in one line: a convincing product name is not evidence.

That line exists because names persuade. Two model names that differ by a letter look like siblings, and a model that has read a thousand catalogs will happily infer that the accessory for one fits the other. Sometimes it does. The harness does not let "sometimes" reach a customer. Eve's answers draw on an exact reference layer (versioned catalog, manual, software, fitment and policy records, with verified resource links) and on a taught library of reviewed lessons found by meaning. An exact claim is either backed by one of those, or it is not made.

Deyaf tells the same story as four steps: listen, look it up, check the rules, reply. The model is one step of four. The other three are plumbing a customer never sees, and they are where most of the honesty is decided.

Courts have already priced the alternative. In February 2024 a British Columbia tribunal decided Moffatt v. Air Canada. The airline's website chatbot had told a grieving customer he could apply for a bereavement fare after he travelled; the airline's own policy page said otherwise. Air Canada argued, in effect, that the chatbot was a separate thing responsible for its own words. The tribunal disagreed and ordered the airline to pay C$812.02. The chatbot was not malicious. It simply answered from somewhere other than the current policy.

Checking the record does not mean handing the model everything. Chroma's study of "context rot" tested 18 models and found that every one degraded as its input grew; on one long-memory benchmark, a focused prompt of about 300 tokens beat the full version of about 113,000 tokens vendor-reported. Anthropic's guide to context engineering makes the same point from the builder's side: good context is the smallest set of high-signal tokens that makes the right outcome likely. The lookup is a narrowing, not a flood. The right record, and only the right record, goes in front of the model.

Everything in the window about 113,000 tokens; the relevant lines are in there somewhere

Only what applies about 300 tokens

Less context, better answers. Schematic and not to scale: the real ratio in Chroma's test was roughly 375 to 1. Source: Chroma, Context Rot, July 2025, where the focused prompt beat the full one (vendor-reported).

Deyaf ↗The model is one step of four

One question, unrolled

Handscrolls were made to be read a little at a time: an arm's width unrolled while the rest waits its turn. Below, one invented question travels through the sequence Eve follows, from the first message to the record it leaves behind.

Illustrative example: not a live Eve transcript. The question, the model names and the records are invented.

  1. A question arrives

    It sounds like one question

    Customer · website chat"Will the mounting kit from my old unit fit the new one?"

    To an honest desk it is several: which old unit, which new one, and what "fit" means here. Nothing has been looked up yet, so nothing can be promised yet.

  2. She listens for what is missing

    One detail, not a form

    Eve"I can check that. Which model is printed on the label of the new unit?"

    Customer"It says 7S."

    She asks for the single detail that unlocks the record. The label beats anyone's memory of a name.

  3. She looks it up

    The exact record, not the likely one

    "7S" sits one letter from "7". A convincing product name is not evidence, so she searches the versioned compatibility records instead of trusting what the names suggest.

    Unit 7 · compatibility recordrev. 2019 Unit 7S · compatibility recordmatch
  4. She checks that it applies

    Real is not the same as relevant

    The current 7S record governs a current 7S. An older manual still explains the original 7, but it cannot set today's options. Authority and applicability outrank whichever page happens to be newest.

  5. If the record answers

    She answers, then stops

    Eve"Yes. The current compatibility record lists your kit for the 7S, and here is its install guide."

    Answer, source, one next step. The reply ends where the evidence ends, not where the model's fluency runs out.

  6. If it does not

    She names the gap

    Eve"I couldn't confirm that pairing in our records, so I won't guess. I can pass it to the team. What's the best way to reach you?"

    A failed lookup means "I couldn't check", not "it doesn't fit". The next step is a hand-off she can actually complete.

  7. What the record keeps

    Nothing lost, nothing invented

    The exchange lands in the conversation record. An unanswered version becomes an unanswered item and, with contact details, a support ticket. She says it is with the team only once the ticket is accepted. A person answers once, and the next customer hears a reviewed lesson.

Scene 1 of 7· scroll, swipe or press ← →

Unknown stays unknown

The second rule is harder, because it asks the system to say less than it could. When sources conflict or a specification is missing, Eve explains the exact uncertainty and offers a useful clarification or a bounded hand-off. A failed lookup means "I couldn't check", not "it doesn't exist". An absence of evidence that two things work together means "unconfirmed", not "incompatible".

The distinction sounds pedantic until it goes wrong. A search that times out and a search that finds nothing can produce the same empty result, and a model that reads both as "no such product" will tell a customer that something they own was never made. "Unknown stays unknown" is the rule that keeps a system failure from turning into a false fact.

The third rule decides which record wins. Authority and applicability outrank recency. The current approved policy governs current purchases; an older manual can explain older equipment, but it cannot set today's options or one owner's eligibility. The newest document in the pile is not automatically the right one, and a page that was never approved does not count at all.

Illustrative · the documents are invented

Sorted by date newest first

  1. Community forum post2026 · the newest page
  2. Compatibility record, 7Scurrent revision · approved
  3. Original manual, Unit 72019 · approved

Sorted by what governs a current 7S authority and applicability first

  1. Compatibility record, 7SGoverns today's options and fit
  2. Original manual, Unit 7Explains the older unit, nothing more
  3. Community forum postNot an approved source; never quoted
Authority and applicability outrank recency. The same three documents, ranked two ways. Only the right-hand order is allowed to shape an answer.

Public failures show what happens when a system forgets this. In April 2025 the AI support agent for the code editor Cursor told users that being logged out when they switched machines was the result of a new policy. There was no such policy; the model had invented a plausible reason, and customers cancelled over it. A year earlier, The Markup found that New York City's MyCity chatbot told business owners they could take a cut of workers' tips and turn away tenants with housing vouchers, both contrary to city law. In each case the system preferred a confident sentence to an honest gap.

Deyaf's harness principles put the lesson bluntly: fake grounding is worse than none. When a source is not connected, the honest reply says so ("that isn't connected yet") and then offers what is known. An invented number feels helpful for one conversation and costs trust for many.

Illustrative example: not a live Eve transcript

The confident invention what a model will say if nothing stops it

"Absolutely! The kit fits every model made since 20191, it's covered by our lifetime warranty2, and if you order today it ships tomorrow3 at 20% off4."

The named uncertainty what is left after every claim is checked

"The current compatibility record lists your kit for the 7S, and here is its install guide. Warranty coverage for your own unit is something the team would need to check."

Check the claims on the left, one at a time, and the honest reply writes itself here.

  1. 1Over-stated. The record confirms one model, the 7S, and says nothing about the rest. She says only what the record covers.
  2. 2Not hers to decide. Coverage for one particular unit is a verdict she does not give, so she names the gap instead.
  3. 3Removed. She never infers stock, freight or delivery dates.
  4. 4Removed. Prices are quoted only as published, with their type. Discounts are never inferred and urgency is never manufactured.
Claims checked: 0 of 4
Editing an invented answer down to what is known. Four claims, four checks: one kept in narrower form, one turned into a named gap, two removed. What remains is shorter, and every word of it can be defended.

Deyaf ↗What happens when she doesn't know

Say only what happened

Words about actions are claims too, and they are the easiest ones to inflate. "I've sent it" sounds better than "I've saved it for the team", so a model eager to please will reach for it. Eve's harness gives her a small, strict vocabulary instead: saved, drafted, sent, uncertain, acknowledged, resolved. Each word names a state that a result has to prove. Saving is not delivery. Delivery is not resolution. Drafting is not sending.

The hand-off is where this matters most. When Eve cannot answer a web, phone, text or email question, it becomes a durable unanswered item, and when contact details are available a support ticket is filed for the team. She says a matter is "with the support team" only when the ticketing system confirms it accepted the ticket. An uncertain submission is not retried blindly, because the other side may already have it; a duplicate ticket is its own small harm.

For this to work, the tools have to tell the truth first. Anthropic's guide to writing tools for agents asks for responses that carry meaningful, high-signal context and for errors that say what actually went wrong instead of an opaque code. A harness can only make an assistant's words as honest as the results underneath them. If the tool says "accepted", she may say "with the team". If it says nothing, she says what she knows.

The discipline shows in what Feniex publishes about hand-offs. From September 16 to 23, 2026, zero were lost: every one sent reached the team. As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional. The hand-off figure on Eve's report card ↗

The same design runs through her other doors. Texts go back to the sender, email replies stay in their thread, and the model never chooses a destination. It can only describe what the system did.

Illustrative scenarios

What does the system actually know?

If her note to the team is stored and no ticket system has replied yet:

The result proves saved

She may say"I've saved your question and your number for the team."

She may not sayThe team has your request.

A save proves the note exists. It proves nothing about anyone else having it.

If a reply is written and waiting for a person to approve it:

The result proves drafted

She may say"I've drafted a reply for the team to review."

She may not sayI've sent you the answer.

Drafting is not sending. Nothing leaves until someone says yes.

If the connection dropped mid-send:

The result proves uncertain

She may say"I'm not certain that went through, so I won't send it twice."

She may not sayDone, it's on its way!

An uncertain outcome is reported as uncertain and never retried blindly: the other side may already have it.

If the message left the system and nobody has confirmed it arrived:

The result proves sent

She may say"I've sent it."

She may not sayThey've got it and they'll call you back.

Sending proves departure. It does not prove arrival, and it certainly does not promise a reply.

If the ticketing system confirms it accepted the ticket:

The result proves acknowledged

She may say"It's with the support team now."

She may not sayIt's fixed, and you'll hear back within the hour.

Acceptance is the proof she needs for "with the team". It is not a fix, and it is not a promise of timing.

If a person answered the question and closed the matter:

The result proves resolved

She may say"The team has resolved this."

She may not sayConsider it resolved on any lower rung.

Resolution belongs to the person who closed the matter. It is the top of the ladder and the only rung where the word is true.

The honest-verb ladder. Pick what the system actually knows; the ladder shows the one verb that result proves, and the sentence it rules out. The dashed rung is the state people forget to name.

Deyaf ↗Answering is not acting

One detail at a time

Restraint also has manners. When a question is missing something, Eve asks for one detail, the one that unlocks the record, and waits. A three-part question feels efficient to the person asking it and tiring to the person answering. Most people answer the part they understand, skip the rest, and the conversation loops back on itself.

Illustrative example: not a live Eve transcript

Three questions at once a generic assistant

Assistant"Which model do you have, when did you buy it, and what exactly isn't fitting?"

Customer"The newer one, I think? It just doesn't fit like the old one did."

One question answered, vaguely. Two skipped. The next turn has to start again.

One detail at a time you choose what Eve asks first

Pick the first detail to ask for. The aim is to reach the exact record in as few turns as possible.

Turns: 0 · Exact record: not found yet

Ask "Which model is on the label?"
"It says 7S." That identifies the exact record, so she can look it up on the next step. One turn.
Ask "When did you buy it?"
"2022, I think." A real detail, but no compatibility record is filed under a purchase year. Still no model.
Ask "What exactly isn't fitting?"
"The holes don't line up." Useful later, but there is nothing to check it against until the model is known.
The detail that unlocks the record comes first. Every other question can wait until there is something exact to compare its answer with.

The same care covers what she does not offer. A price is quoted only as published and with its type. She never infers discounts, stock or delivery dates, never manufactures urgency, and claims no authority to approve returns, change accounts, place orders or decide sensitive matters. When someone describes a hazard, safety comes before troubleshooting: stop, step away, get help, and no walkthrough into opening anything up.

These limits are not timidity. In December 2023 a car dealership's website chatbot was talked into "agreeing" to sell a new SUV for one dollar and calling it a binding offer. The lesson the industry took from that screenshot was structural: a model should never make a commitment that code has not approved. Deyaf publishes a three-level ladder (answer only, prepare for approval, run an approved task), set per workflow and per channel, and every assistant starts on the first level. There is no single switch that lets her do everything.

Deyaf ↗What the assistant can and cannot do

The box that waits

Everything in this essay eventually has to fit inside a small box in the corner of a web page. The website chat box is the most visible door a harness has, and the smallest, and it is where restraint is easiest to lose: a box that can speak is tempted to speak first.

So the manners begin before the first word. A restrained box waits to be opened; it does not pop up over the page someone came to read. When it does speak, it says what it is. Eve introduces herself as Feniex's intelligence and never pretends to be a person. Her reply streams, so the first words arrive quickly, and the page the visitor is on travels with the question, so she does not have to ask what they are looking at. Then the shape from the first chapter: the answer, one useful next step, and a stop, usually in two or three sentences. The box goes quiet and leaves the next move to the visitor.

The last courtesy is an honest path to a person. When she cannot answer, the question becomes a durable unanswered item; with contact details a ticket is filed, and she says the matter is with the support team only once that ticket is confirmed accepted. The Box in the Corner draws that hand-off panel by panel, beside the stream and the product tile, and From Script to Source tells how the box got here.

On the phone the pause means something else. Silence on a call sounds like a dropped line, so there the harness fills it with short cues, "One moment" after about a second and a half. Restraint on a call is not quiet; it is brief. Nobody Has to Hold follows that line through a day.

Deyaf ↗Watch Eve answer on five doors

A desk that is taught

There are two broad ways to make an assistant better over time. One lets the agent improve itself: it writes its own skills, curates its own memory and grows around the person using it. The other keeps a human in the loop for everything that becomes public knowledge. Eve is built the second way, on purpose. She is a specialist front desk, and a stranger's first five minutes are no place for a self-taught fact.

In practice this looks like a quiet queue. Every question she could not answer lands as an unanswered item. A person writes the reusable answer once; it is reviewed; the lesson and its receipt are committed together; and the next customer who asks gets it. Answering this customer and creating a reusable lesson are separate operations. Raw conversations never become public knowledge automatically, and contact information is stored apart from published knowledge and never copied into it.

At a larger scale, support tickets and calls are folded into a redacted, de-duplicated question book. Eve is measured on a frozen test set without making live changes; only explicitly approved lessons are published; then she is measured again. Deyaf lists the lesson sources as support tickets (monthly), call recordings (every 60 days) and the teach queue (daily). It is a supervised knowledge-improvement loop: not fine-tuning, and not training on every raw conversation.

The engineering community has been arriving at the same instinct from another direction. Mitchell Hashimoto, describing how he adopted coding agents, calls it engineering the harness: when an agent makes a mistake, change its instructions or its tools so that it does not make that mistake again. The team behind the Manus agent argues for keeping failed steps in the context instead of scrubbing them, because erasing a failure erases the evidence a model needs to adapt. Eve's version is slower and more supervised. A miss is not hidden; it is queued, answered once by a person, and measured before it is taught.

The difference matters most on the days nobody is watching. A self-improving generalist can drift toward whatever its last hundred conversations rewarded. A taught desk can only say what a person has written or approved, and its answers cite the approved source they came from. Humans author public knowledge; the harness measures and proves it.

Deyaf ↗The loop that makes her better

Measured, not promised

Restraint is only a virtue if it is measured, because a system that refuses everything is never technically wrong. Eve's report card counts restraint as a result, not as a failure to answer. On 100 fixed test questions graded on September 23, 2026, 86% were handled well, and 10% contained a material error. Inside the 86: 65 fully answered, 17 handled safely within limits, and 4 where she asked the right question back. Another 4 were useful but partial. There were no critical errors.

  • 0 critical errors

86% handled well (65 + 17 + 4), shown beside 10% material error.

As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional. Graded September 23, 2026 on 100 fixed test questions. Not an independent benchmark. See the report card on Deyaf ↗

The trend is reported honestly too: 88% on September 17, 90% on September 21, 86% on September 23. That is not steady improvement, and nobody presents it as such. By type, reworded questions scored 100% and manual questions 95%, while past trouble spots scored 40%. That last group is kept in the test on purpose, so the weakest ground stays in view. Anthropic's guide to agent evaluations recommends building a first test set from a few dozen real failures; keeping the hard cases in is the same idea, applied to a front desk.

The numbers are Feniex's own, machine-judged and provisional, and they are published with their weak spots showing. That is the last form of restraint in the harness: claiming no more about the system than the system has shown.

Deyaf ↗Eve's report card, weak spots included

Build yours

Eve is Feniex's assistant. Deyaf is the product that lets your business build its own. Yours starts from the same kind of harness: knowledge you approve, doors you choose, rules you set and a team that teaches it.

The builder takes about ten minutes to a knowledge preview. Name the assistant, add what it should know, and watch it quote your own sources, with a hand-off when it does not know. It is a preview, not live AI, and live doors switch on one at a time.

An ink-painted calligraphy brush lies at rest beside a small pool of ink on pale paper.

Questions, answered briefly

Each answer here follows the rule this essay describes: the answer, one next step, then a stop.

Is an Agentic Harness just a longer prompt?
No. A prompt is one input to the model; the harness is everything around it: the records it may quote, the memory it keeps, the doors it answers, the actions it may take and the people who teach it. The definition in chapter one is the place to start.
Why not let the model answer from what it already knows?
Because fluency is not evidence, and general knowledge cannot tell you which revision of a policy is current. Exact claims come from the applicable record, or they are not made.
Doesn't all this restraint make an assistant less helpful?
It changes what helpful means: a named gap and a hand-off that actually lands beat a confident guess. Eve's report card counts "handled safely, within limits" as a good outcome.
Does Eve learn from every conversation?
No. Lessons are written or approved by people, measured on a frozen test set, and published only when approved.
Is Eve a person?
No. Eve is software with a named personality, with no age, body or experiences; there is more in the FAQ at the bottom of the Meet Eve page ↗
Is the Deyaf builder the same thing as Eve?
No. The builder is a knowledge preview that quotes the notes you give it, in your browser; it is not live AI, and it is not Eve's voice. Live doors are switched on one at a time, after activation.

Sources, shelved

What this essay leans on, and for what. These sources informed the writing; none of them endorses Deyaf or is affiliated with it. Eve's figures come from Feniex's own dated measurements, labelled where they appear.

  1. Building Effective AI AgentsAnthropic Engineering · December 2024Used for: start with the simplest solution; add complexity only when it helps.
  2. Effective context engineering for AI agentsAnthropic Engineering · September 2025Used for: the smallest set of high-signal context.
  3. Writing effective tools for agentsAnthropic Engineering · September 2025Used for: tool results and errors that say what really happened.
  4. Demystifying evals for AI agentsAnthropic Engineering · January 2026Used for: building test sets from real failures.
  5. Context RotChroma Research · July 2025 · vendor-reportedUsed for: 18 models, focused (about 300 tokens) versus full (about 113,000) prompts.
  6. My AI Adoption JourneyMitchell Hashimoto · February 2026Used for: engineering the harness after every mistake.
  7. Context Engineering for AI Agents: Lessons from Building ManusManus · July 2025Used for: keeping failures visible instead of erasing them.
  8. Moffatt v. Air CanadaMcCarthy Tétrault TechLex · February 2024 · on a misrepresentation by an AI chatbotUsed for: answering from the current policy source.
  9. AI Incident Database, Incident 1039April 2025 · Cursor's support agent invents a policyUsed for: never inventing policy.
  10. NYC's AI chatbot tells businesses to break the lawThe Markup · March 29, 2024Used for: authoritative grounding over confident sentences.
  11. AI Incident Database, Incident 622December 2023 · a dealership chatbot agrees to a one-dollar SUVUsed for: no commitments that code has not approved.

Colophon

Set in Shippori Mincho and Zen Kaku Gothic New. Paper white #F7F5EF, ink #1E1C1A, indigo #2C3E57, moss #8A8F5C and stone #B9B2A5; vermilion appears only on the seals and as a single accent in the ink paintings; each seal opens a page on deyaf.com.

The ink paintings were made for this page with an image model and art-directed to carry mood only. Every diagram and interactive figure is drawn by hand in HTML and SVG. Conversations are invented and labelled as illustrations. Published 25 September 2026 by Deyaf by Feniex ↗