Harness Library Preprints No. 15 · v1 Build yours ↗

Agentic HarnessHarness Library Preprints · No. 15

PreprintEvaluationVersion 1Not peer reviewed

Measured, Not PromisedHow to evaluate a customer-facing assistant, and why the harness around the model is what makes it measurable

Figure 1 · Graphical abstract a, b schematic · c measured

aWhere lessons come fromLesson sources, in deyaf.com's own wording

Lesson sources, in deyaf.com's wording. Support tickets (monthly) and call recordings (every 60 days) are folded into a redacted, de-duplicated question book. The teach queue (daily) brings questions to a person, who answers once; the answer is checked, then taught. Separately, the teach loop measures on its own frozen test set, with no live changes. Support ticketsmonthly Call recordingsevery 60 days Teach queuedaily Question book redacted, de-duplicated Answered once by a person; checked, then taught The teach loop's frozen test set measured with no live changes · see b

bHow a change is allowed inMeasured before she's taught, then measured again

A round starts by measuring on the frozen set with no live changes. An automatic stop halts the run if a measurement tried to write, grading is incomplete, or the first batch fails its proof. Otherwise only approved lessons are taught, each with a dated receipt, and the assistant is measured again on taught questions and untaught controls. Then the next round begins. 1 · Measure firstfrozen set, no live changes Any stop? yes: halt before the next write no 2 · Teach approved lessonslesson + receipt, approval dated 3 · Measure againtaught questions + untaught controls next round

cWhat the last measurement saidEve, 100 fixed test questions, graded 23 Sep 2026

86%handled well
10%material error, shown at equal weight
  • 65 fully answered
  • 17 handled safely, within limits
  • 4 asked the right question back
  • 4 useful but partial
  • 10 material error
  • 0 critical errors

Source: Eve's report card on agenticharness.com. As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional.

Figure 1 | The paper in one picture. a, Lesson sources. Tickets and calls are folded into a redacted, de-duplicated question book; the teach queue brings questions to a person, who answers once. The teach loop measures on its own frozen test set. b, Every change passes the same gate: measure with no live changes, stop automatically if anything is off, teach only approved lessons, then measure again against untaught controls. c, The most recent public measurement of Eve, Feniex's assistant. The report card is its own fixed test of 100 questions, separate from the teach loop in a and b; each square is one of those questions; material errors are drawn in the darkest colour and marked so they cannot hide inside the headline. Panels a and b are schematics; panel c is measured data.

Abstract

A customer-facing assistant is easy to demonstrate and hard to measure. This paper explains what an Agentic Harness is and argues that the harness is also what makes an assistant measurable: it fixes the questions, records what happened, and separates answering from acting. We review four evaluation ideas that rarely reach marketing copy: passk, which asks whether an assistant succeeds every time rather than once; capability versus regression suites; code, model and human graders; and the counting rule hidden inside any "resolution rate". Interactive figures let the reader change the inputs, including an ablation explorer that removes harness parts one at a time. As a case study we describe how Eve, Feniex's customer assistant, is measured before she is taught, and reproduce her latest public report card in full: 86% of 100 fixed test questions handled well, with a 10% material-error rate beside it (Feniex's own measurements, Sep 17–24, 2026).

Keywords · Agentic Harness · evaluation · passk · regression suite · graders · customer service · human hand-off

A grayscale, electron-microscope-style view of a regular honeycomb lattice, sharp in the foreground and softening into the distance.
Plate I. A regular lattice, seen only because something was measured. Generated illustration; it carries no data.

1Introduction

Every assistant demo succeeds. The question a business has to answer is different: of the next thousand customers, how many get a right answer, how many get a safe hand-off, and how many get something wrong, said with confidence? The model alone cannot answer that, because the model is not what the customer meets.

At Feniex, customers meet Eve. She answers on the website, by voice, on the phone, by text and by email, and she helps people choose equipment, decide on a purchase, or get support for what they already own. When she does not know, she says so and hands the question to a person. Behind her is not one large model but a system around one, and that system is our subject.

Definition 1 · Agentic HarnessAn Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.

Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. For the whole system chapter by chapter, see how an Agentic Harness works.

Our claim is modest and, we think, under-appreciated: the harness is also the measuring instrument. A bare model produces text. A harness produces a record: which question arrived, which source was checked, which rule applied, what was said and what happened next. Without that record there is nothing to grade. With it, a business can ask the question that matters before it widens an assistant's scope: on questions nobody chose after the fact, how often is it right, how often is it safe, and how often is it wrong?

Section 2 locates measurement inside a single answer. Section 3 reviews methods. Section 4 lets you remove harness parts and watch the numbers move. Sections 5 and 6 describe how Eve is measured before she is taught, and report her latest results, weak spots included. Section 7 splits the measurement by the two doors customers meet most, the chat box and the phone.

2What the harness adds, and where measurement attaches

LangChain's shorthand is a useful start: an agent is a model plus a harness [11]. The model predicts text. The harness decides what the model sees, what it may call, what it may say, and what is kept afterwards. Anthropic describes the working rhythm of an agent as gather context, take action, verify the work, repeat [13]. Verification is the step most often left out.

deyaf.com breaks one customer answer into four steps: listen, look it up, check the rules, reply. The model is one step of four. Figure 2 draws those steps as a trace, the way observability tools record agent work as nested spans [9], and marks where each step can be measured.

Birgitta Böckeler's taxonomy sorts the probes [10]. Some parts of a harness are guides: they shape behaviour before it happens, like a rulebook or an approved library. Others are sensors: they observe what happened and report back, like a check that a cited record exists or a grader that compares an answer with the team's. Either kind can be computational, meaning deterministic code, or inferential, meaning a model does the judging. Computational sensors are cheap and exact but narrow. Inferential sensors can judge completeness, but they inherit a model's errors, which is why their verdicts are provisional.

Figure 2 · One answer as a trace, with its sensors Illustrative trace · two measured marks
Span
Sensor · kind
One turn website chat, whole answer
Outcome graded against the team's answer model
Listen read the question, detect the language
Language detected correctly code
Look it up exact records: catalog, manuals, policy
Cited record exists and is current code
Look it up taught library, by meaning
Lesson is a reviewed one code
Check the rules is this action allowed here?
Claim matches what the result proves code
Reply the model writes; words stream
Complete, on-scope, right tone model person
Figure 2 | Where measurement attaches to one answer. The four steps are deyaf.com's own (the model is one step of four); span lengths are illustrative. Only the two vertical marks are measured: in Eve's website chat the first word arrives at 1.4 s and the reply is finished at 2.7 s (Feniex's own measurements, Sep 17–24, 2026; see Eve's numbers, measured not promised, and her dated report card). The sensor column shows the kind of check a harness can attach to each span; most are cheap code checks, and only judgements of quality need a model or a person. The sensors are examples of what a harness can attach, not a list of Eve's checks.

A customer-facing harness adds one property that makes measurement tractable: it separates answering from acting. Eve's public actions are few (look something up, search the taught library, take a message, record an unanswered question) and she describes only what a result proves. Saving is not delivery, delivery is not resolution, and drafting is not sending. Because the claims are narrow, they can be checked. An assistant that may do anything can only be evaluated by watching everything.

3Methods: evaluating a customer-facing assistant

3.1Start from failures, then freeze the set

Anthropic's evaluation guide recommends starting with a few dozen tasks drawn from real failures rather than an exhaustive benchmark; 20 to 50 is enough to begin [1]. The same guide separates capability evaluations, which ask what an assistant can do and are expected to start low, from regression evaluations, which ask whether it still does what it used to and should sit near 100%.

Two disciplines follow. First, freeze the set: if the questions change between runs, a rising score may only mean easier questions. Second, keep the embarrassing questions. A regression suite that quietly drops the items an assistant once failed drifts toward flattery. Eve's report card does the opposite and keeps a group of past trouble spots on purpose; that group scores 40%, against 86% handled well overall with a 10% material-error rate (Section 6; Feniex's own measurements, Sep 17–24, 2026).

In practice these suites belong inside the harness, not in a notebook. deyaf.com states the discipline in four words, measured before she's taught, and describes the loop that makes her better. At Feniex, every release is also checked at every door and in every language with test turns that cannot write anything, and a person reads the verdict.

3.2One success is not reliability: pass@k and passk

Most benchmarks report whether an assistant can solve a task. A customer cares whether it will, this time and next time. Two statistics capture the difference [1]. If one attempt succeeds with probability p, and attempts are independent:

pass@k = 1 − (1 − p)k(1)

passk = pk(2)

pass@k rewards retries and climbs toward certainty; passk punishes inconsistency and falls. At p = 0.90, pass@8 is effectively 100% while pass8 is 43%. Sierra's τ-bench made this concrete for customer-service agents: strong models solved under half of its retail tasks, and roughly a quarter when the same task had to succeed eight times (vendor-reported) [2].

A harness changes the arithmetic by turning failures into safe outcomes. If a check catches most wrong answers and converts them into a hand-off, the chance that a conversation does no harm rises well above the chance that it answers, and passk for harm-free conversations decays slowly. Security writing makes the same point more bluntly: Simon Willison argues that a defence which works 95% of the time is a failing one [14].

Figure 3 · Reliability over repeated tries Illustrative model

Static view: p = 0.90, checks catch 90% of failures, k = 8. The sliders need JavaScript; the source data below covers the same curves.

At k = 8: pass@8 >99.9% · pass8 43.0% · with checks 92.3%. Eight tries in a row is roughly one busy morning of the same question.

Source data · p = 0.90, checks catch 90%
kpass@kpasskwith checks
190.0%90.0%99.0%
299.0%81.0%98.0%
4>99.9%65.6%96.1%
8>99.9%43.0%92.3%
16>99.9%18.5%85.1%
Figure 3 | The same assistant, three ways of counting its tries. Equations (1) and (2) with independent attempts. "With checks" assumes a harness check catches the stated share of failures and turns each into a safe hand-off, so the per-try chance of doing no harm is p + (1 − p) × catch rate. The numbers are an illustrative model, not measurements of any assistant. Drag k upward and watch pass@k and passk diverge.

3.3Who grades: code, models and people

Graders come in three kinds: code, models and people [1]. Code checks are exact (did the answer cite a record that exists, was the reply in the customer's language?) but cannot judge whether an explanation was complete. Model graders scale and can compare an answer with a reference, but they are themselves probabilistic. People are the reference standard and the most expensive.

Field data suggests teams know this. In a study of agents running in production, 74% relied on human evaluation, and 68% handed to a human within ten steps [3]. LangChain's 2026 survey of 1,340 practitioners found quality to be the top barrier to production (32%), yet only 37.3% evaluated live traffic, against 89% with some form of observability [4]. Watching is common; grading is not. The honest label for a model-graded score is the one Feniex uses: machine-judged and provisional. It is a measurement, not a certification.

3.4What counts as "resolved"

No customer-service metric is quoted more, or defined less consistently, than the resolution rate. Intercom's documentation separates a resolution the customer confirms from an assumed one, where the customer simply stops replying for 24 hours (vendor definitions) [6]. Zendesk announced pricing per automated resolution, so the definition carries money [7]. Salesforce's own Customer Zero pages have reported an 83% resolution figure for January 2025 and a 63% figure in 2026 (vendor-reported) [5]; before comparing two such numbers, you would need to know exactly what each one counted.

Figure 4 shows why. The same 100 synthetic conversations are scored under three rules. Assumed resolution counts silence as success, including seven wrong answers whose customers never came back, and scores 74%. Confirmed resolution counts only what customers confirmed and scores 51%, missing sixteen right answers nobody acknowledged. A graded rule checks every conversation against a known right answer, counts safe hand-offs as handled, and reports 83% with an 11% material-error rate beside it.

Figure 4 · Same 100 conversations, three scores Synthetic data
  • Right answer: 46 confirmed by the customer, 16 went quiet, 5 after one clarifying question
  • Handed to a person: 12 resolved by them, 4 still open
  • Wrong answer: 7 customers went quiet, 4 came back
  • Abandoned before any answer: 6
  • Outlined: not counted as resolved by this rule
Counting rule
74%"resolved"

Counts 7 wrong answers as wins because those customers stopped replying, and drops all 16 hand-offs, even the 12 that a person resolved.

51%resolved

Misses 16 right answers nobody confirmed, and says nothing at all about the 11 wrong ones.

83%handled well11%material error

Needs a known right answer for every question and a grader you trust, and reports the error rate beside the headline. (In this illustration, the 6 abandoned conversations count as not handled.)

Figure 4 | A resolution rate is a counting rule with a number attached. One hundred synthetic conversations, fixed in advance; only the rule changes. Filled dots are counted as resolved, outlined dots are not, and a cross marks a wrong answer. Switching rules moves the headline by 32 points without a single conversation changing. The data are invented for illustration and describe no real assistant.

None of these rules is dishonest in itself. What misleads is the flattering rule reported without its definition, or a success rate reported without an error rate beside it. That is why every Eve figure in this paper carries its date, its source and its material-error rate.

4Ablations: what each part of the harness buys

An ablation removes one component and measures the difference. It is how you show that a part matters instead of asserting it. Harness work has produced a few public examples: LangChain kept the model fixed, changed only the harness, and moved a coding benchmark from 52.8 to 66.5 (vendor-reported) [8]. Customer service has a sharper cautionary tale. In Moffatt v. Air Canada, decided on 14 February 2024, a tribunal held an airline responsible for what its website chatbot told a customer about bereavement fares [15], which is the kind of error that answering only from the current policy record is designed to prevent.

Figure 5 is an ablation explorer built on an illustrative model. Its numbers are invented to show direction, not size, and they are not Eve's. Remove any of four parts (grounding in exact records, the human gate, the teach loop and the regression suite) and watch five readings move.

Figure 5 · Ablation explorer Illustrative model · not Eve's data

Static view: the full harness. With JavaScript, untick a part to remove it and the readings update.

Customer question
ModelWrites the words. Always on: agent = model + harness.
Reply
Around the loop
Dashboard "resolved"what a naive report shows
64%baseline
Graded correct and safechecked against the right answer
82%baseline
Material errorsper 100 conversations
5baseline
Incidentsunapproved promises reaching a customer, per 100
0.2baseline
Median time to a finished replyseconds
4.0 sbaseline

Full harness: 82% of conversations graded correct and safe, 5 material errors and 0.2 incidents per 100. The dashboard shows only 64% "resolved", because hand-offs are not counted as resolutions. That undercount is the price of honesty.

Figure 5 | Removing a check can raise apparent speed and lower trust. Each reading shows the full harness (open circle), the current configuration (dot) and an uncertainty band. Without the human gate nothing is handed off, so the dashboard's "resolved" rises and replies finish sooner, while graded quality falls and incidents climb. Without the regression suite the central estimate barely moves but the band widens: the assistant is no worse on the day; you simply stop knowing when it gets worse. Illustrative model: effects are additive with one interaction (grounding and gate both removed), invented for teaching, and not measurements of Eve or any product.

Two patterns are worth naming. First, removing a check often improves the numbers a dashboard shows; the cost appears only in graded quality and in incidents, meaning promises the business never approved. Second, removing the regression suite hides problems rather than creating them. That is why Deyaf's public ladder has three levels (answer only, prepare for approval, run an approved task) and why each task moves up separately: answering is not acting. Autonomy is earned in measured stages.

5Case study: how Eve is measured before she is taught

Eve is Feniex's warm, composed public face. She answers first, gives one useful next step, then stops. She asks for one missing detail at a time, and she checks the applicable record before any exact product, compatibility, warranty, price or software claim, because a convincing product name is not evidence. She speaks English, Spanish, French, Portuguese and Arabic from one English library. On the phone she answers Feniex's overflow and after-hours line when no one on the team is free, with no transfers, by design: she answers from her library, or takes the caller's name and number and hands the matter to the team.

For an evaluator, the interesting part is the loop that decides what she knows. deyaf.com puts it in one line, your team answers once; she knows it for good, and Figure 6 walks through one round.

5.1Where the questions come from

deyaf.com lists the lesson sources as support tickets (monthly), call recordings (every 60 days) and the teach queue (daily). Tickets and calls are folded into a redacted, de-duplicated question book. When Eve cannot answer a question on any door, it becomes a durable unanswered item. Through the teach queue a person answers it once; the answer is checked, then taught, and next time she knows.

5.2Answering and teaching are separate operations

Answering a customer and creating a reusable lesson are different operations. A person answers once; the answer is checked, then taught. In Teach Eve, the approved question-and-answer pair is embedded, and the lesson and its receipt are committed together. Approval is stamped as a date, never a name, and a lesson can be quarantined and later restored. Raw conversations never become public knowledge automatically.

5.3Measure first, with nothing written

Before lessons are published, Eve is measured on a frozen test set without making live changes. The run has automatic stops: it halts if a measurement tried to write, if grading is incomplete, or if the first batch fails its proof. Only explicitly approved lessons are published. After teaching she is measured again, and a set of untaught control questions is measured alongside the taught ones. The taught questions should move; the controls should not. If the controls move too, something other than the lessons changed, and the round is not evidence of learning.

A grayscale, electron-microscope-style view of crystal growth terraces rising in fine, even steps.
Plate II. Growth in steps, each layer laid on a checked one. Generated illustration; it carries no data.
Figure 6 · One round of measure, teach, re-measure Illustrative round

Static view: the whole round, ending in the verdict. With JavaScript, step through it and try a failed proof.

  1. Queue.

    Unanswered questions become durable items in the teach queue.

    Illustrative example: not a live Eve transcript
    1. "Does the warranty still apply if a dealer installed it?" asked 7×
    2. "Which manual covers the older model?" 4×
    3. "Can I update the software myself?" 3×
  2. A person answers once.

    The answer is checked before it goes any further.

  3. Measure first.

    Frozen test set, no live changes. Baseline recorded for taught-to-be and control questions.

  4. Teach, then prove.

    Only approved lessons; lesson and receipt committed together, approval dated. The first batch must pass its proof or the run halts.

  5. Measure again.

    Same frozen set, taught questions and untaught controls side by side.

  6. Read the verdict.

    Taught questions moved, controls held, any regression flagged. A lesson can be quarantined and restored.

Verdict: taught questions +44 points; controls −1 point, within noise. One control question regressed and is flagged for review.

Figure 6 | Measured before taught, then measured again with controls. The steps follow the loop Feniex describes publicly; the numbers are an illustrative round, not Eve's results. The untaught controls are what make the chart evidence: a rise in taught questions means little unless questions nobody touched stay where they were. With JavaScript on, tick "fail the first batch's proof" to see an automatic stop halt the run before anything else is written.

5.4What the loop is not

This is a supervised knowledge-improvement loop, not fine-tuning, and not training on every raw conversation. The model's weights do not change; the library does, one reviewed lesson at a time. One of the harness's design rules says it plainly: humans author public knowledge, and the harness measures and proves it.

6Results: Eve's report card

Table 1 reproduces the report card deyaf.com publishes, graded on 23 September 2026 on 100 fixed test questions. Of the 100 answers, 65 were fully answered, 17 were handled safely within limits, 4 asked the right question back, 4 were useful but partial, and 10 contained a material error. There were no critical errors. The headline, 86% handled well, is always reported with the 10% material-error rate beside it.

Table 1 · Eve's report card Measured · Feniex, graded 23 Sep 2026
a · How all 100 answers landed
OutcomenShare
Fully answered65
Handled safely, within limits17
Asked the right question back4
Useful but partial4
Material error10
Critical error0
Handled well86%
Material-error rate10%
b · Handled well, by question type
Question typeScoreofBar
Reworded questions100%15/15
From the manuals95%38/40
Trick questions95%19/20
Real customer wording80%8/10
Past trouble spots†40%6/15
Handled-well trend: 88 percent on 17 September, 90 percent on 21 September, 86 percent on 23 September. 88%90%86% 17 Sep21 Sep23 Sep

c · Trend, handled well. 88% → 90% → 86%. Not a steady rise: the latest point is the lowest of the three. On 100 fixed questions, a four-point move is four questions; with a machine grader, that is within run-to-run variation, so read the three points as neither progress nor a trend.

†Past trouble spots are questions she used to get wrong, kept in the test on purpose. Per-type scores sum to the 86 handled well.

Table 1 | Graded on a fixed test, weak spots included. Reproduced from Eve's report card (weak spots included). The success and error rates are set at the same type size on purpose.As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional.

Three details matter more than the headline. The trend runs 88%, 90%, 86%: not a steady rise, and the latest point is the lowest. By question type, reworded questions scored 100% and manual-based questions 95%, but real customer wording scored 80%, and past trouble spots scored 40%, all beside the same 10% material-error rate (Feniex's own measurements, Sep 17–24, 2026). That last group is exactly where the next lessons should come from, which is the loop in Section 5 working as intended.

Speed is measured separately: about three seconds to a finished reply in website chat and on the phone. The knowledge behind the answers, counted on 24 September 2026, includes 558 company documents and 2,000+ taught lessons (Feniex's own measurements, Sep 17–24, 2026; see Eve's report card, weak spots included). None of these figures is an independent benchmark; all are Feniex's own measurements of its own assistant.

7Measuring the two doors: the chat box and the phone

The report card grades answers, not doors. Its 100 fixed questions are not split by channel, so the headline, 86% handled well beside a 10% material-error rate, is best read as one number for one library, whichever way a question arrives. Eve keeps the same identity, library and rules at every door. What the doors change is the rest of the measurement: how long a person waits, and what "handled" has to mean when nobody is looking at a screen.

Time. Speed is measured per door (Table 2a). In website chat the first word arrives at 1.4 s and the finished reply at 2.7 s; on the phone, 1.6 s and 3.0 s. A caller has no typing indicator, so the phone line is designed never to go silent: "One moment" after about 1.5 s, "Let me check on that" when a lookup starts, "Still checking" about every 4 s. Overflow and after-hours call taking has its own study in the library, Nobody Has to Hold.

Coverage. The phone adds a count the chat box does not have (Table 2b). The team's phones ring first; when no one is free or the office is closed, Eve answers instead of voicemail. There are no transfers, by design: she answers from her library, or takes the caller's name and number and hands the matter to the team. So the right numerator is not "calls resolved" but calls answered, callers who stayed, and hand-offs with the details taken. For the line itself, see Eve on the phone.

Table 2 · The two doors, measured Measured · Feniex, Sep 17–24, 2026
a · Time to answer, by door
DoorFirst wordFinishedBar
Website chat1.4 s2.7 s
Phone line1.6 s3.0 s
Handled well86%
Material-error rate10%
b · The phone line, first week (since 17 Sep 2026)
Countn
Calls answered102
… while the team was busy86
… after the office closed16
Callers who stayed and talked93%
Handed straight to the team (details taken, never a transfer)60

Bars in a: dark segment to the first word, light segment to the finished reply, on one scale. The 86% and 10% come from the 100-question report card (Table 1) and are not split by door.

Table 2 | Two doors, one library. Speed is about three seconds to a finished reply at both doors; the phone line also counts who got answered. Figures from Eve's report card, with the error rate beside the score.As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional.

8Discussion and limitations

A report card is an instrument, and instruments have limits. We note five.

Sample size. One hundred questions is a small set. On 100 fixed questions, a four-point move is four questions; nothing is re-sampled between runs, so the variation comes from the model and the machine grader, and a move that size is within run-to-run variation. Read the three points as neither progress nor a trend. The grader. The scores are machine-judged. A model grader can be wrong in both directions, which is why the label says provisional and why human review remains the reference. Representativeness. A frozen set trades realism for comparability; the 80% on real customer wording is the closest signal to live traffic, and it sits below the 86% headline, which travels with a 10% material-error rate (Feniex's own measurements, Sep 17–24, 2026).

Not a benchmark. These are Feniex's measurements of Feniex's assistant. They say nothing about another business's assistant, which starts fresh with its own knowledge and has to be measured on its own questions. What is not measured. The report card grades answers. It does not measure satisfaction, and it cannot see questions nobody asked.

A grayscale, electron-microscope-style field of identical packed spheres with a single sphere missing and a faint dislocation line.
Plate III. A perfect array with one site missing. A measurement worth trusting keeps its defect in view. Generated illustration; it carries no data.

Against these limits sits one strength that is rarer than it should be. The error is published with the success, the weak spots are kept rather than trimmed, and a lower number was reported when it came. Deyaf is built from Eve: the harness that runs Eve at Feniex, packaged so another business can have an assistant of its own. What it inherits is less a score than a habit. Measured, not promised.

AGlossary

pass@k
The chance that at least one of k attempts succeeds. Flattering for retries.
passk
The chance that all k attempts succeed. The reliability a customer actually experiences.
Capability suite
Tasks the assistant is expected to find hard; scores start low and should climb.
Regression suite
Tasks it used to pass; scores should stay near 100%, and a drop is a finding.
Frozen test set
Questions fixed between runs so that two scores can be compared.
Material error
An answer wrong in a way that matters to the customer. Always reported beside the success rate.
Untaught controls
Questions no lesson touched, measured to catch side effects of teaching.
Ablation
Removing one component to measure what it contributes.
Grader
The code, model or person that decides whether an answer passed.

BBoundaries the measurements assume

Measurement means something only inside boundaries. Eve claims no authority to approve returns, change accounts, place orders or decide sensitive matters, and she says a matter is "with the support team" only when the provider confirms it accepted the ticket. Deyaf's starting plan requires human review before external messages, record changes or commitments. Deyaf claims no security certifications; live customer service requires activation and a completed security review. See what the assistant can and cannot do.

End matter

Reproducibility

You can run a miniature of this loop on your own knowledge. Deyaf's builder takes about ten minutes, needs no account, and ends in a knowledge preview that quotes your notes, with no live AI. Ask it something your notes don't cover and it shows the hand-off; teach the answer and the preview uses it next time. Try the builder: about ten minutes, no account.

Data availability

The report-card figures are published by Deyaf at agenticharness.com/eve, section 05. Every other number in this paper comes from an illustrative model or synthetic data, labelled as such, or from the cited sources.

Acknowledgements

To the Feniex team, whose answers become lessons, and to Eve, Feniex's assistant, who is software with a named personality, not a person, and did not review this paper.

Competing interests

This paper is published by Deyaf, which packages the harness Eve runs on. Cited sources are independent and do not endorse Deyaf. Vendor figures are reported as their publishers stated them.

ROpen review

An editorial device, not a real peer review: no external review took place. These are the questions readers ask most, set out as reviewer comments with the authors' responses.

R1.1Reviewer 1

Is 86% good?

ResponseIt depends on the other number. With a 10% material-error rate, about one answer in ten said something wrong that mattered. Whether that is acceptable depends on the stakes and on whether errors are caught before they reach a customer, which is why the two numbers travel together (Feniex's own measurements, Sep 17–24, 2026; the report card).

R1.2Reviewer 1

Why publish a score that went down?

ResponseA report card that only ever rises is a press release. The trend is 88%, 90%, 86% handled well, with a 10% material-error rate beside the latest (Feniex's own measurements, Sep 17–24, 2026; see the trend). Readers should see the dip, and a frozen set makes it comparable.

R1.3Reviewer 1

Keeping questions she used to fail looks like self-sabotage.

ResponseIt is the regression suite doing its job. Past trouble spots scored 40%, beside an overall 86% handled well and a 10% material-error rate (Feniex's own measurements, Sep 17–24, 2026; by question type). Trimming them would lift the headline and hide exactly the questions the next round of teaching should target.

R2.1Reviewer 2

Isn't a model grading a model circular?

ResponsePartly, which is why the result is labelled machine-judged and provisional. Model graders scale; people remain the reference. The honest move is to say which one produced the number.

R2.2Reviewer 2

Does Eve learn from every conversation?

ResponseNo. A person answers once, the lesson is checked and approved, and she is measured before and after. The model is not fine-tuned, and raw conversations never become public knowledge automatically.

R2.3Reviewer 2

Would an assistant built for my business score the same?

ResponseNobody can say before measuring. An assistant built with Deyaf starts fresh with your knowledge, and its report card has to be earned on your own questions.

R3.1Reviewer 3

What is an Agentic Harness, in one sentence?

ResponseEverything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it. Eve is the voice; the harness is everything behind her.

Measure your own assistant from day one

Start with a name, a little knowledge and one good job. Watch where it hands off, teach the answer, and keep the error rate beside the headline.

Build your assistant
Or start with the team-knowledge template

References

  1. 1Anthropic. Demystifying evals for AI agents. Engineering blog, January 2026. anthropic.com/engineering/demystifying-evals-for-ai-agents
  2. 2Sierra. τ-bench: shaping development and evaluation of agents. June 2024. vendor-reported sierra.ai/blog/tau-bench-shaping-development-evaluation-agents
  3. 3Measuring Agents in Production. arXiv preprint 2512.04123. arxiv.org/abs/2512.04123
  4. 4LangChain. State of Agent Engineering. 12 June 2026, n = 1,340. langchain.com/state-of-agent-engineering
  5. 5Salesforce. Agentforce: Customer Zero. vendor-reported salesforce.com/news/stories/agentforce-customer-zero
  6. 6Intercom. Fin AI Agent resolutions. Help centre. vendor definitions fin.ai/help/en/articles/10772642-fin-ai-agent-resolutions
  7. 7Zendesk. Zendesk introduces outcome-based pricing. 28 August 2024. zendesk.com/newsroom/articles/zendesk-outcome-based-pricing
  8. 8LangChain. Improving Deep Agents with harness engineering. 17 February 2026. vendor-reported langchain.com/blog/improving-deep-agents-with-harness-engineering
  9. 9OpenTelemetry. GenAI observability. 2026. opentelemetry.io/blog/2026/genai-observability
  10. 10B. Böckeler. Harness engineering for coding agent users. martinfowler.com, 2 April 2026. martinfowler.com/articles/harness-engineering.html
  11. 11LangChain. The Anatomy of an Agent Harness. 10 March 2026. langchain.com/blog/the-anatomy-of-an-agent-harness
  12. 12M. Hashimoto. My AI Adoption Journey. 5 February 2026. mitchellh.com/writing/my-ai-adoption-journey
  13. 13Anthropic. Building agents with the Claude Agent SDK. September 2025. claude.com/blog/building-agents-with-the-claude-agent-sdk
  14. 14S. Willison. The lethal trifecta for AI agents. 16 June 2025. simonwillison.net/2025/Jun/16/the-lethal-trifecta
  15. 15McCarthy Tétrault. Moffatt v. Air Canada: a misrepresentation by an AI chatbot. Decision of 14 February 2024. mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot

Cited sources are independent of Deyaf and do not endorse it. Tool and vendor names appear only as industry sources, never as a description of Eve's stack.

Cite this article

@article{ahbl2026measured,
  title   = {Measured, Not Promised: How to evaluate a customer-facing
             assistant, and why the harness makes it measurable},
  author  = {{Deyaf Editorial}},
  journal = {Agentic Harness Blog Library},
  number  = {15},
  year    = {2026},
  month   = sep,
  note    = {Preprint, version 1. Case study: Eve at Feniex.}
}