Abstract
A customer-facing assistant is easy to demonstrate and hard to measure. This paper explains what an Agentic Harness is and argues that the harness is also what makes an assistant measurable: it fixes the questions, records what happened, and separates answering from acting. We review four evaluation ideas that rarely reach marketing copy: passk, which asks whether an assistant succeeds every time rather than once; capability versus regression suites; code, model and human graders; and the counting rule hidden inside any "resolution rate". Interactive figures let the reader change the inputs, including an ablation explorer that removes harness parts one at a time. As a case study we describe how Eve, Feniex's customer assistant, is measured before she is taught, and reproduce her latest public report card in full: 86% of 100 fixed test questions handled well, with a 10% material-error rate beside it (Feniex's own measurements, Sep 17–24, 2026).
Keywords · Agentic Harness · evaluation · passk · regression suite · graders · customer service · human hand-off
1Introduction
Every assistant demo succeeds. The question a business has to answer is different: of the next thousand customers, how many get a right answer, how many get a safe hand-off, and how many get something wrong, said with confidence? The model alone cannot answer that, because the model is not what the customer meets.
At Feniex, customers meet Eve. She answers on the website, by voice, on the phone, by text and by email, and she helps people choose equipment, decide on a purchase, or get support for what they already own. When she does not know, she says so and hands the question to a person. Behind her is not one large model but a system around one, and that system is our subject.
Definition 1 · Agentic HarnessAn Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.
Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. For the whole system chapter by chapter, see how an Agentic Harness works.
Our claim is modest and, we think, under-appreciated: the harness is also the measuring instrument. A bare model produces text. A harness produces a record: which question arrived, which source was checked, which rule applied, what was said and what happened next. Without that record there is nothing to grade. With it, a business can ask the question that matters before it widens an assistant's scope: on questions nobody chose after the fact, how often is it right, how often is it safe, and how often is it wrong?
Section 2 locates measurement inside a single answer. Section 3 reviews methods. Section 4 lets you remove harness parts and watch the numbers move. Sections 5 and 6 describe how Eve is measured before she is taught, and report her latest results, weak spots included. Section 7 splits the measurement by the two doors customers meet most, the chat box and the phone.
2What the harness adds, and where measurement attaches
LangChain's shorthand is a useful start: an agent is a model plus a harness [11]. The model predicts text. The harness decides what the model sees, what it may call, what it may say, and what is kept afterwards. Anthropic describes the working rhythm of an agent as gather context, take action, verify the work, repeat [13]. Verification is the step most often left out.
deyaf.com breaks one customer answer into four steps: listen, look it up, check the rules, reply. The model is one step of four. Figure 2 draws those steps as a trace, the way observability tools record agent work as nested spans [9], and marks where each step can be measured.
Birgitta Böckeler's taxonomy sorts the probes [10]. Some parts of a harness are guides: they shape behaviour before it happens, like a rulebook or an approved library. Others are sensors: they observe what happened and report back, like a check that a cited record exists or a grader that compares an answer with the team's. Either kind can be computational, meaning deterministic code, or inferential, meaning a model does the judging. Computational sensors are cheap and exact but narrow. Inferential sensors can judge completeness, but they inherit a model's errors, which is why their verdicts are provisional.
A customer-facing harness adds one property that makes measurement tractable: it separates answering from acting. Eve's public actions are few (look something up, search the taught library, take a message, record an unanswered question) and she describes only what a result proves. Saving is not delivery, delivery is not resolution, and drafting is not sending. Because the claims are narrow, they can be checked. An assistant that may do anything can only be evaluated by watching everything.
3Methods: evaluating a customer-facing assistant
3.1Start from failures, then freeze the set
Anthropic's evaluation guide recommends starting with a few dozen tasks drawn from real failures rather than an exhaustive benchmark; 20 to 50 is enough to begin [1]. The same guide separates capability evaluations, which ask what an assistant can do and are expected to start low, from regression evaluations, which ask whether it still does what it used to and should sit near 100%.
Two disciplines follow. First, freeze the set: if the questions change between runs, a rising score may only mean easier questions. Second, keep the embarrassing questions. A regression suite that quietly drops the items an assistant once failed drifts toward flattery. Eve's report card does the opposite and keeps a group of past trouble spots on purpose; that group scores 40%, against 86% handled well overall with a 10% material-error rate (Section 6; Feniex's own measurements, Sep 17–24, 2026).
In practice these suites belong inside the harness, not in a notebook. deyaf.com states the discipline in four words, measured before she's taught, and describes the loop that makes her better. At Feniex, every release is also checked at every door and in every language with test turns that cannot write anything, and a person reads the verdict.
3.2One success is not reliability: pass@k and passk
Most benchmarks report whether an assistant can solve a task. A customer cares whether it will, this time and next time. Two statistics capture the difference [1]. If one attempt succeeds with probability p, and attempts are independent:
pass@k = 1 − (1 − p)k(1)
passk = pk(2)
pass@k rewards retries and climbs toward certainty; passk punishes inconsistency and falls. At p = 0.90, pass@8 is effectively 100% while pass8 is 43%. Sierra's τ-bench made this concrete for customer-service agents: strong models solved under half of its retail tasks, and roughly a quarter when the same task had to succeed eight times (vendor-reported) [2].
A harness changes the arithmetic by turning failures into safe outcomes. If a check catches most wrong answers and converts them into a hand-off, the chance that a conversation does no harm rises well above the chance that it answers, and passk for harm-free conversations decays slowly. Security writing makes the same point more bluntly: Simon Willison argues that a defence which works 95% of the time is a failing one [14].
Static view: p = 0.90, checks catch 90% of failures, k = 8. The sliders need JavaScript; the source data below covers the same curves.
At k = 8: pass@8 >99.9% · pass8 43.0% · with checks 92.3%. Eight tries in a row is roughly one busy morning of the same question.
| k | pass@k | passk | with checks |
|---|---|---|---|
| 1 | 90.0% | 90.0% | 99.0% |
| 2 | 99.0% | 81.0% | 98.0% |
| 4 | >99.9% | 65.6% | 96.1% |
| 8 | >99.9% | 43.0% | 92.3% |
| 16 | >99.9% | 18.5% | 85.1% |
3.3Who grades: code, models and people
Graders come in three kinds: code, models and people [1]. Code checks are exact (did the answer cite a record that exists, was the reply in the customer's language?) but cannot judge whether an explanation was complete. Model graders scale and can compare an answer with a reference, but they are themselves probabilistic. People are the reference standard and the most expensive.
Field data suggests teams know this. In a study of agents running in production, 74% relied on human evaluation, and 68% handed to a human within ten steps [3]. LangChain's 2026 survey of 1,340 practitioners found quality to be the top barrier to production (32%), yet only 37.3% evaluated live traffic, against 89% with some form of observability [4]. Watching is common; grading is not. The honest label for a model-graded score is the one Feniex uses: machine-judged and provisional. It is a measurement, not a certification.
3.4What counts as "resolved"
No customer-service metric is quoted more, or defined less consistently, than the resolution rate. Intercom's documentation separates a resolution the customer confirms from an assumed one, where the customer simply stops replying for 24 hours (vendor definitions) [6]. Zendesk announced pricing per automated resolution, so the definition carries money [7]. Salesforce's own Customer Zero pages have reported an 83% resolution figure for January 2025 and a 63% figure in 2026 (vendor-reported) [5]; before comparing two such numbers, you would need to know exactly what each one counted.
Figure 4 shows why. The same 100 synthetic conversations are scored under three rules. Assumed resolution counts silence as success, including seven wrong answers whose customers never came back, and scores 74%. Confirmed resolution counts only what customers confirmed and scores 51%, missing sixteen right answers nobody acknowledged. A graded rule checks every conversation against a known right answer, counts safe hand-offs as handled, and reports 83% with an 11% material-error rate beside it.
- Right answer: 46 confirmed by the customer, 16 went quiet, 5 after one clarifying question
- Handed to a person: 12 resolved by them, 4 still open
- Wrong answer: 7 customers went quiet, 4 came back
- Abandoned before any answer: 6
- Outlined: not counted as resolved by this rule
Counts 7 wrong answers as wins because those customers stopped replying, and drops all 16 hand-offs, even the 12 that a person resolved.
Misses 16 right answers nobody confirmed, and says nothing at all about the 11 wrong ones.
Needs a known right answer for every question and a grader you trust, and reports the error rate beside the headline. (In this illustration, the 6 abandoned conversations count as not handled.)
None of these rules is dishonest in itself. What misleads is the flattering rule reported without its definition, or a success rate reported without an error rate beside it. That is why every Eve figure in this paper carries its date, its source and its material-error rate.
4Ablations: what each part of the harness buys
An ablation removes one component and measures the difference. It is how you show that a part matters instead of asserting it. Harness work has produced a few public examples: LangChain kept the model fixed, changed only the harness, and moved a coding benchmark from 52.8 to 66.5 (vendor-reported) [8]. Customer service has a sharper cautionary tale. In Moffatt v. Air Canada, decided on 14 February 2024, a tribunal held an airline responsible for what its website chatbot told a customer about bereavement fares [15], which is the kind of error that answering only from the current policy record is designed to prevent.
Figure 5 is an ablation explorer built on an illustrative model. Its numbers are invented to show direction, not size, and they are not Eve's. Remove any of four parts (grounding in exact records, the human gate, the teach loop and the regression suite) and watch five readings move.
Static view: the full harness. With JavaScript, untick a part to remove it and the readings update.
Full harness: 82% of conversations graded correct and safe, 5 material errors and 0.2 incidents per 100. The dashboard shows only 64% "resolved", because hand-offs are not counted as resolutions. That undercount is the price of honesty.
Two patterns are worth naming. First, removing a check often improves the numbers a dashboard shows; the cost appears only in graded quality and in incidents, meaning promises the business never approved. Second, removing the regression suite hides problems rather than creating them. That is why Deyaf's public ladder has three levels (answer only, prepare for approval, run an approved task) and why each task moves up separately: answering is not acting. Autonomy is earned in measured stages.
5Case study: how Eve is measured before she is taught
Eve is Feniex's warm, composed public face. She answers first, gives one useful next step, then stops. She asks for one missing detail at a time, and she checks the applicable record before any exact product, compatibility, warranty, price or software claim, because a convincing product name is not evidence. She speaks English, Spanish, French, Portuguese and Arabic from one English library. On the phone she answers Feniex's overflow and after-hours line when no one on the team is free, with no transfers, by design: she answers from her library, or takes the caller's name and number and hands the matter to the team.
For an evaluator, the interesting part is the loop that decides what she knows. deyaf.com puts it in one line, your team answers once; she knows it for good, and Figure 6 walks through one round.
5.1Where the questions come from
deyaf.com lists the lesson sources as support tickets (monthly), call recordings (every 60 days) and the teach queue (daily). Tickets and calls are folded into a redacted, de-duplicated question book. When Eve cannot answer a question on any door, it becomes a durable unanswered item. Through the teach queue a person answers it once; the answer is checked, then taught, and next time she knows.
5.2Answering and teaching are separate operations
Answering a customer and creating a reusable lesson are different operations. A person answers once; the answer is checked, then taught. In Teach Eve, the approved question-and-answer pair is embedded, and the lesson and its receipt are committed together. Approval is stamped as a date, never a name, and a lesson can be quarantined and later restored. Raw conversations never become public knowledge automatically.
5.3Measure first, with nothing written
Before lessons are published, Eve is measured on a frozen test set without making live changes. The run has automatic stops: it halts if a measurement tried to write, if grading is incomplete, or if the first batch fails its proof. Only explicitly approved lessons are published. After teaching she is measured again, and a set of untaught control questions is measured alongside the taught ones. The taught questions should move; the controls should not. If the controls move too, something other than the lessons changed, and the round is not evidence of learning.
Static view: the whole round, ending in the verdict. With JavaScript, step through it and try a failed proof.
- Queue.
Unanswered questions become durable items in the teach queue.
Illustrative example: not a live Eve transcript- "Does the warranty still apply if a dealer installed it?" asked 7×
- "Which manual covers the older model?" 4×
- "Can I update the software myself?" 3×
- A person answers once.
The answer is checked before it goes any further.
- Measure first.
Frozen test set, no live changes. Baseline recorded for taught-to-be and control questions.
- Teach, then prove.
Only approved lessons; lesson and receipt committed together, approval dated. The first batch must pass its proof or the run halts.
- Measure again.
Same frozen set, taught questions and untaught controls side by side.
- Read the verdict.
Taught questions moved, controls held, any regression flagged. A lesson can be quarantined and restored.
Verdict: taught questions +44 points; controls −1 point, within noise. One control question regressed and is flagged for review.
5.4What the loop is not
This is a supervised knowledge-improvement loop, not fine-tuning, and not training on every raw conversation. The model's weights do not change; the library does, one reviewed lesson at a time. One of the harness's design rules says it plainly: humans author public knowledge, and the harness measures and proves it.
6Results: Eve's report card
Table 1 reproduces the report card deyaf.com publishes, graded on 23 September 2026 on 100 fixed test questions. Of the 100 answers, 65 were fully answered, 17 were handled safely within limits, 4 asked the right question back, 4 were useful but partial, and 10 contained a material error. There were no critical errors. The headline, 86% handled well, is always reported with the 10% material-error rate beside it.
| Outcome | n | Share |
|---|---|---|
| Fully answered | 65 | |
| Handled safely, within limits | 17 | |
| Asked the right question back | 4 | |
| Useful but partial | 4 | |
| Material error | 10 | |
| Critical error | 0 | |
| Handled well | 86% | |
| Material-error rate | 10% |
| Question type | Score | of | Bar |
|---|---|---|---|
| Reworded questions | 100% | 15/15 | |
| From the manuals | 95% | 38/40 | |
| Trick questions | 95% | 19/20 | |
| Real customer wording | 80% | 8/10 | |
| Past trouble spots† | 40% | 6/15 |
c · Trend, handled well. 88% → 90% → 86%. Not a steady rise: the latest point is the lowest of the three. On 100 fixed questions, a four-point move is four questions; with a machine grader, that is within run-to-run variation, so read the three points as neither progress nor a trend.
†Past trouble spots are questions she used to get wrong, kept in the test on purpose. Per-type scores sum to the 86 handled well.
Three details matter more than the headline. The trend runs 88%, 90%, 86%: not a steady rise, and the latest point is the lowest. By question type, reworded questions scored 100% and manual-based questions 95%, but real customer wording scored 80%, and past trouble spots scored 40%, all beside the same 10% material-error rate (Feniex's own measurements, Sep 17–24, 2026). That last group is exactly where the next lessons should come from, which is the loop in Section 5 working as intended.
Speed is measured separately: about three seconds to a finished reply in website chat and on the phone. The knowledge behind the answers, counted on 24 September 2026, includes 558 company documents and 2,000+ taught lessons (Feniex's own measurements, Sep 17–24, 2026; see Eve's report card, weak spots included). None of these figures is an independent benchmark; all are Feniex's own measurements of its own assistant.
7Measuring the two doors: the chat box and the phone
The report card grades answers, not doors. Its 100 fixed questions are not split by channel, so the headline, 86% handled well beside a 10% material-error rate, is best read as one number for one library, whichever way a question arrives. Eve keeps the same identity, library and rules at every door. What the doors change is the rest of the measurement: how long a person waits, and what "handled" has to mean when nobody is looking at a screen.
Time. Speed is measured per door (Table 2a). In website chat the first word arrives at 1.4 s and the finished reply at 2.7 s; on the phone, 1.6 s and 3.0 s. A caller has no typing indicator, so the phone line is designed never to go silent: "One moment" after about 1.5 s, "Let me check on that" when a lookup starts, "Still checking" about every 4 s. Overflow and after-hours call taking has its own study in the library, Nobody Has to Hold.
Coverage. The phone adds a count the chat box does not have (Table 2b). The team's phones ring first; when no one is free or the office is closed, Eve answers instead of voicemail. There are no transfers, by design: she answers from her library, or takes the caller's name and number and hands the matter to the team. So the right numerator is not "calls resolved" but calls answered, callers who stayed, and hand-offs with the details taken. For the line itself, see Eve on the phone.
| Door | First word | Finished | Bar |
|---|---|---|---|
| Website chat | 1.4 s | 2.7 s | |
| Phone line | 1.6 s | 3.0 s | |
| Handled well | 86% | ||
| Material-error rate | 10% | ||
| Count | n |
|---|---|
| Calls answered | 102 |
| … while the team was busy | 86 |
| … after the office closed | 16 |
| Callers who stayed and talked | 93% |
| Handed straight to the team (details taken, never a transfer) | 60 |
Bars in a: dark segment to the first word, light segment to the finished reply, on one scale. The 86% and 10% come from the 100-question report card (Table 1) and are not split by door.
8Discussion and limitations
A report card is an instrument, and instruments have limits. We note five.
Sample size. One hundred questions is a small set. On 100 fixed questions, a four-point move is four questions; nothing is re-sampled between runs, so the variation comes from the model and the machine grader, and a move that size is within run-to-run variation. Read the three points as neither progress nor a trend. The grader. The scores are machine-judged. A model grader can be wrong in both directions, which is why the label says provisional and why human review remains the reference. Representativeness. A frozen set trades realism for comparability; the 80% on real customer wording is the closest signal to live traffic, and it sits below the 86% headline, which travels with a 10% material-error rate (Feniex's own measurements, Sep 17–24, 2026).
Not a benchmark. These are Feniex's measurements of Feniex's assistant. They say nothing about another business's assistant, which starts fresh with its own knowledge and has to be measured on its own questions. What is not measured. The report card grades answers. It does not measure satisfaction, and it cannot see questions nobody asked.
Against these limits sits one strength that is rarer than it should be. The error is published with the success, the weak spots are kept rather than trimmed, and a lower number was reported when it came. Deyaf is built from Eve: the harness that runs Eve at Feniex, packaged so another business can have an assistant of its own. What it inherits is less a score than a habit. Measured, not promised.
AGlossary
- pass@k
- The chance that at least one of k attempts succeeds. Flattering for retries.
- passk
- The chance that all k attempts succeed. The reliability a customer actually experiences.
- Capability suite
- Tasks the assistant is expected to find hard; scores start low and should climb.
- Regression suite
- Tasks it used to pass; scores should stay near 100%, and a drop is a finding.
- Frozen test set
- Questions fixed between runs so that two scores can be compared.
- Material error
- An answer wrong in a way that matters to the customer. Always reported beside the success rate.
- Untaught controls
- Questions no lesson touched, measured to catch side effects of teaching.
- Ablation
- Removing one component to measure what it contributes.
- Grader
- The code, model or person that decides whether an answer passed.
BBoundaries the measurements assume
Measurement means something only inside boundaries. Eve claims no authority to approve returns, change accounts, place orders or decide sensitive matters, and she says a matter is "with the support team" only when the provider confirms it accepted the ticket. Deyaf's starting plan requires human review before external messages, record changes or commitments. Deyaf claims no security certifications; live customer service requires activation and a completed security review. See what the assistant can and cannot do.
End matter
Reproducibility
You can run a miniature of this loop on your own knowledge. Deyaf's builder takes about ten minutes, needs no account, and ends in a knowledge preview that quotes your notes, with no live AI. Ask it something your notes don't cover and it shows the hand-off; teach the answer and the preview uses it next time. Try the builder: about ten minutes, no account.
Data availability
The report-card figures are published by Deyaf at agenticharness.com/eve, section 05. Every other number in this paper comes from an illustrative model or synthetic data, labelled as such, or from the cited sources.
Acknowledgements
To the Feniex team, whose answers become lessons, and to Eve, Feniex's assistant, who is software with a named personality, not a person, and did not review this paper.
Competing interests
This paper is published by Deyaf, which packages the harness Eve runs on. Cited sources are independent and do not endorse Deyaf. Vendor figures are reported as their publishers stated them.
ROpen review
An editorial device, not a real peer review: no external review took place. These are the questions readers ask most, set out as reviewer comments with the authors' responses.
R1.1Reviewer 1
Is 86% good?
ResponseIt depends on the other number. With a 10% material-error rate, about one answer in ten said something wrong that mattered. Whether that is acceptable depends on the stakes and on whether errors are caught before they reach a customer, which is why the two numbers travel together (Feniex's own measurements, Sep 17–24, 2026; the report card).
R1.2Reviewer 1
Why publish a score that went down?
ResponseA report card that only ever rises is a press release. The trend is 88%, 90%, 86% handled well, with a 10% material-error rate beside the latest (Feniex's own measurements, Sep 17–24, 2026; see the trend). Readers should see the dip, and a frozen set makes it comparable.
R1.3Reviewer 1
Keeping questions she used to fail looks like self-sabotage.
ResponseIt is the regression suite doing its job. Past trouble spots scored 40%, beside an overall 86% handled well and a 10% material-error rate (Feniex's own measurements, Sep 17–24, 2026; by question type). Trimming them would lift the headline and hide exactly the questions the next round of teaching should target.
R2.1Reviewer 2
Isn't a model grading a model circular?
ResponsePartly, which is why the result is labelled machine-judged and provisional. Model graders scale; people remain the reference. The honest move is to say which one produced the number.
R2.2Reviewer 2
Does Eve learn from every conversation?
ResponseNo. A person answers once, the lesson is checked and approved, and she is measured before and after. The model is not fine-tuned, and raw conversations never become public knowledge automatically.
R2.3Reviewer 2
Would an assistant built for my business score the same?
ResponseNobody can say before measuring. An assistant built with Deyaf starts fresh with your knowledge, and its report card has to be earned on your own questions.
R3.1Reviewer 3
What is an Agentic Harness, in one sentence?
ResponseEverything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it. Eve is the voice; the harness is everything behind her.
Measure your own assistant from day one
Start with a name, a little knowledge and one good job. Watch where it hands off, teach the answer, and keep the error rate beside the headline.
References
- 1Anthropic. Demystifying evals for AI agents. Engineering blog, January 2026. anthropic.com/engineering/demystifying-evals-for-ai-agents
- 2Sierra. τ-bench: shaping development and evaluation of agents. June 2024. vendor-reported sierra.ai/blog/tau-bench-shaping-development-evaluation-agents
- 3Measuring Agents in Production. arXiv preprint 2512.04123. arxiv.org/abs/2512.04123
- 4LangChain. State of Agent Engineering. 12 June 2026, n = 1,340. langchain.com/state-of-agent-engineering
- 5Salesforce. Agentforce: Customer Zero. vendor-reported salesforce.com/news/stories/agentforce-customer-zero
- 6Intercom. Fin AI Agent resolutions. Help centre. vendor definitions fin.ai/help/en/articles/10772642-fin-ai-agent-resolutions
- 7Zendesk. Zendesk introduces outcome-based pricing. 28 August 2024. zendesk.com/newsroom/articles/zendesk-outcome-based-pricing
- 8LangChain. Improving Deep Agents with harness engineering. 17 February 2026. vendor-reported langchain.com/blog/improving-deep-agents-with-harness-engineering
- 9OpenTelemetry. GenAI observability. 2026. opentelemetry.io/blog/2026/genai-observability
- 10B. Böckeler. Harness engineering for coding agent users. martinfowler.com, 2 April 2026. martinfowler.com/articles/harness-engineering.html
- 11LangChain. The Anatomy of an Agent Harness. 10 March 2026. langchain.com/blog/the-anatomy-of-an-agent-harness
- 12M. Hashimoto. My AI Adoption Journey. 5 February 2026. mitchellh.com/writing/my-ai-adoption-journey
- 13Anthropic. Building agents with the Claude Agent SDK. September 2025. claude.com/blog/building-agents-with-the-claude-agent-sdk
- 14S. Willison. The lethal trifecta for AI agents. 16 June 2025. simonwillison.net/2025/Jun/16/the-lethal-trifecta
- 15McCarthy Tétrault. Moffatt v. Air Canada: a misrepresentation by an AI chatbot. Decision of 14 February 2024. mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot
Cited sources are independent of Deyaf and do not endorse it. Tool and vendor names appear only as industry sources, never as a description of Eve's stack.
Cite this article
@article{ahbl2026measured,
title = {Measured, Not Promised: How to evaluate a customer-facing
assistant, and why the harness makes it measurable},
author = {{Deyaf Editorial}},
journal = {Agentic Harness Blog Library},
number = {15},
year = {2026},
month = sep,
note = {Preprint, version 1. Case study: Eve at Feniex.}
}