Build note 07 · Agentic Harness engineering · 11 min read

Reliability needs receipts.

Subject Agentic Harness

A model can produce an answer. An Agentic Harness must be able to show how that answer became an action: which context entered, which tools ran, which permissions applied, what was checked, and who approved the consequence.

An original Deyaf perspective informed by Emily Winks’ Atlan build guide, published April 13 and updated June 4, 2026. Atlan does not endorse or partner with Deyaf.

Chain-of-custody recordRUN / 007
01Trusted contextsource identified
02Bounded toolscope checked
03Permission gatepolicy enforced
04Verificationoutput tested
05Human decisionimpact approved

The operating thesis

The harness is not just what helps an agent act. It is what makes the action inspectable, bounded, and reversible.

Chapter

01

Start before the first prompt.

Atlan’s most useful provocation is its “Step 0”: certify the data layer before building the reasoning loop. That reframes reliability. A valid query against the wrong table is still a wrong answer.

An Agentic Harness is the operating layer around a model. It supplies instructions and context, exposes tools, holds state, enforces permissions, controls execution, records traces, validates results, and routes consequential actions through human oversight. The model reasons; the harness decides what the reasoning is allowed to touch and what must happen before its output becomes work.

Deyaf’s own definition is plainer: “An Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.” See how an Agentic Harness works, end to end.

That means the first design question is not “Which framework?” It is “What information is trusted enough to enter the run?” Certification, glossary definitions, ownership, freshness, and lineage form the beginning of the chain of custody. They let the system distinguish an available asset from an approved one.

Visual 01 · What changes when context carries governanceIllustrative mechanics — not benchmark output

Bare schema

Available is mistaken for trusted.

The model sees names and types, but not which asset is current, certified, or meaningful to the business.

QUESTION
“What was recognized revenue?”

VISIBLE
revenue_v1
revenue_final
finance_rev_new
?Table selected by resemblancemeaning unknown
!Query succeedsanswer untrusted

Governed context

Trust travels with the asset.

The harness resolves a business term, checks certification and lineage, then passes the approved context into the run.

TERM
recognized_revenue

RESOLVES TO
finance.revenue_booked
certified · owner: finance
1Definition resolvedone meaning
2Asset certifiedpolicy passes
3Lineage attachedorigin visible
38%

Atlan reports a 38% improvement in AI-generated SQL accuracy when rich semantic metadata was supplied in its testing. Its public report page describes 522 evaluations across 174 queries. This is Atlan’s own study, not an independent industry-wide benchmark. Review Atlan’s methodology summary and download page ↗

Chapter

02

Build the custody stack.

A framework may supply orchestration primitives. A production harness is the assembled operating system: each layer answers a different question about trust.

The build order matters because later controls depend on earlier evidence. Permissions cannot evaluate a resource whose identity is unclear. Observability cannot explain a run whose state was never recorded. Evals cannot isolate regressions if model and harness changes are mixed together.

Select a layer below to inspect its role. The “done test” is intentionally operational: a layer exists only when its behavior can be demonstrated, not when a configuration file has been created.

Visual 02 · Six records in the custody stackInteractive layer inspection

Record 01 · Input custody

Context says what the agent is allowed to know.

Retrieve only the material relevant to this step, attach provenance and definitions, and exclude deprecated or unapproved sources before the model reasons.

DONE TEST
The same request resolves to the same governed definition, and every supplied asset exposes its source and status.IN EVE
Her exact reference layer holds versioned catalog, manual, software, fitment and policy records with verified resource links, and she checks it before any exact claim.

Chapter

03

The loop needs an exit.

ReAct formalized an interleaving of reasoning and action. Production engineering adds the pieces a research pattern does not promise: step ceilings, deterministic checks, permission gates, and a deliberate stopping rule.

The original ReAct paper describes a useful rhythm: reason about the task, act through an external interface, observe the result, and update the plan. The harness turns that rhythm into controlled execution. It counts steps, records tool results, detects failure, and decides whether the next operation is permitted.

Human intervention belongs where mistakes are hard to reverse. OpenAI’s practical guide recommends human oversight for high-risk or irreversible actions and when the agent repeatedly exceeds failure thresholds. The point is not to put a person in every loop. It is to put judgment at the boundary where impact changes.

At Feniex, Eve’s loop is deliberately short: she listens, looks it up, checks the rules, then replies. The model is one step of four. Look inside one of Eve’s answers.

Visual 03 · A bounded execution loopText summary follows the diagram
Diagram summary: inference proposes; policy permits; the event record preserves; verification tests; the stop condition ends. If the tool would create material impact, the path pauses for a person before execution.

Chapter

04

Treat traces as product evidence.

A trace is more than a debugging artifact. It is the receipt that lets a team reconstruct what happened, discover new failure modes, and turn those failures into regression tests.

Persist state as events: request accepted, context resolved, tool proposed, permission evaluated, result observed, verification passed, approval recorded. This creates restartability and accountability at once. A stopped run can resume from a known state; a completed run can be reviewed without relying on the model’s summary of its own behavior.

The strongest eval sets grow from real traces. When production reveals a failure the suite missed, that trace becomes a test case. Run those cases when the harness changes, and score them independently so the system is not grading its own work.

Publishing the weak spots is part of the receipt. Deyaf shows Eve’s report card with its trouble spots left in, and the next chapter opens her case file.

Visual 04 · Case record 01: one run, reconstructed from eventsIllustrative event log
Certified context resolved3 assets
Read tool permittedpolicy r-12
Result stored in task stateevent 018
Write requires approvalheld
Human approval recordedreviewer 02
Verification suite passed6 / 6

Harness experiment

+13.7 pts

LangChain reports moving a coding agent from 52.8% to 66.5% on Terminal-Bench 2.0 while keeping the model fixed and changing the harness. This is a vendor-published result on one benchmark, not a universal gain. Read the experiment ↗

Durable instructions

AGENTS.md

The open format gives coding agents a predictable place to find setup, tests, conventions, and safety notes. The project says it is used by more than 60,000 open-source projects. Review the standard ↗

Chapter

05

Now inspect a working desk.

Every record above is illustrative. This one has a name. Eve is Feniex’s assistant, and Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop.

Plate · Case file 02, open on the deskMood illustration — carries no data
Illustration: seen from above on an ink-dark desk, five blank cream slips are clipped into a chain across an open cobalt folder. A coiled phone cord feeds the first slip, an acid-green tab marks the last, and a thin red thread runs from the fourth slip to a separate blank card.

She is Feniex’s warm, composed public face, and she introduces herself as “Eve, Feniex’s intelligence.” She helps people choose suitable equipment, make a buying decision, or get support for equipment they own. She is software with a named personality, not a person, and approved Feniex operators manage her rules and knowledge. Meet Eve, Feniex’s assistant.

Her context arrives with custody attached. Before any exact product, compatibility, warranty, price or software claim, she checks the applicable record: versioned catalog, manual, software, fitment and policy records, then a library of reviewed lessons. A convincing product name is not evidence, and unknown stays unknown.

The same Eve answers at every door: website chat, website voice, phone, text and email. On the phone she is the overflow and after-hours line; the team’s phones ring first. Her memory is split into four kinds, each with one job. Customer notes are short and historical, never current facts: orders, invoices, shipments and balances are always looked up fresh from the live account record.

Visual 05 · Case record 02: one question, one hand-offIllustrative example: not a live Eve transcript
Question arrives by phone, in Spanishdoor · phone
Rendered into English for the one libraryone library
Exact record checked before the claimreference layer
Answer given in Spanish, from that recordsource kept
Caller asks for a personno transfers, by design
Name and callback number takenmessage
Ticket provider confirms it was acceptedconfirmed
Only now: “with the support team”claim = result

Her action surface is small on purpose: four public actions, to look something up, search the taught library, take a message, and record an unanswered question. She claims no authority to approve returns, change accounts or place orders. On Deyaf’s published ladder, answering is not acting.

When she doesn’t know, the question becomes a durable unanswered item, and a support ticket is filed when contact details are available. That is what happens when she doesn’t know, and it doubles as a regression source: tickets and calls are folded into a redacted, de-duplicated question book, she is measured on a frozen test set without live changes, only approved lessons are published, and she is measured again. A supervised loop, not training on every raw conversation.

Report card · graded Sep 23, 2026

86%

“Handled well” on 100 fixed test questions, shown beside a 10% material-error rate and 0 critical errors. Past trouble spots scored 40%; they stay in the test on purpose. Eve’s report card, weak spots included ↗

As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional.

Hand-offs · Sep 16–23, 2026

0 lost

Every hand-off sent reached the team. A hand-off passes the customer’s details to people; it is never a live transfer.

As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional.

Chapter

06

Two doors, two receipts.

Most people meet the harness in one of two places: a chat box on a web page, or a call the team could not pick up. Both keep the custody rules above, and both leave a record.

The chat box is the smallest door and the most visible. Page and product context ride along, the reply streams in, and product or resource tiles can arrive with it. She answers first, offers one useful next step, then stops, usually within two or three sentences. She asks for one missing detail at a time and never pretends to be a person. Step inside the box, drawn as a comic, read how chat went from scripts to sources, or check the manners a chat box owes its visitors.

On the phone, the team’s phones ring first. When no one is free or the office is closed, Eve answers instead of voicemail: “Hello, this is Eve, Feniex’s intelligence. How can I help you?” She is never silent for long: “One moment” after about 1.5 seconds, “Let me check on that” as a lookup starts, “Still checking” about every four seconds. No transfers, by design: she answers from her library, or takes a name and number and hands the matter to the team. Read why nobody has to hold, why “press 1” is over, and the parts of a phone assistant, as a periodic table.

Receipt rule · chat box

Proved

She says only what the result proves. An unanswered question becomes a durable item; with contact details a support ticket is filed, and “with the support team” appears only once it is confirmed accepted. One assistant, every way in ↗

Receipt rule · phone line

Transcribed

Every call is transcribed. A recorded fallback plays if she is down, calls run ten minutes at most, and she never reads web addresses aloud. From Sep 17, 2026, her first week: 102 calls answered, 86 while the team was busy, 16 after closing. The dated figures, error rate included ↗

As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional.

Chapter

07

Scale the controls, not the ambiguity.

The same custody logic applies across organization sizes. What changes is the number of workflows, owners, identities, environments, and review obligations that the harness must coordinate.

Decision aid · Change the organization profileIllustrative planning guidance

Illustrative profile · Small team

One workflow. One owner. One visible boundary.

Start narrow enough that a person can understand every tool, data source, and approval rule. The goal is not maximum autonomy; it is a useful task with an observable failure envelope.

  • Choose one repeatable workflow with a measurable finish
  • Expose only the few tools required for that workflow
  • Require approval before external sends or material writes
  • Review every failed run and add one regression test

Deyaf’s opportunity is to make this operating pattern accessible without pretending every organization needs the same stack. Deyaf is built from Eve, Feniex’s assistant, who already answers customers at every door. Deyaf’s early-access builder lets your business set up its own; live doors switch on one at a time.

Larger teams should read the boundaries before the features. Deyaf’s trust and control principles claim no security certifications or compliance guarantees: live customer service requires activation and a completed security review.

01

Scope

Can the job and its successful finish fit in one clear sentence?

02

Context

Can every input expose its source, owner, freshness, and status?

03

Authority

Are reads, writes, and irreversible actions classified differently?

04

Evidence

Can a reviewer replay the run from recorded events?

05

Recovery

Can the system stop, resume, reverse, or fail without hidden impact?

The next record starts here

Build the proof around the intelligence.

The model supplies capability. The harness supplies custody: trusted context, bounded tools, visible state, enforceable permissions, verification, and accountable approval.

Deyaf’s builder is an early-access knowledge preview: about ten minutes to a saved setup, no account needed.

Build your assistant