Architecture note · Reliability beyond the model

The system around intelligence.

The Agentic Harness.

A language model can reason. An agent can act. But only an Agentic Harness can reliably connect what the model knows, what the business allows, what the tools can do, and how the result gets checked.

The propositionModel quality ≠ system reliability
The unit of designThe complete operating loop
The control surfaceContext, tools, policy, proof
Reading time11 minutes
The signal
Many failures blamed on the model are actually failures of the architecture around it.
01The definition

An Agentic Harness is the operational layer.

It is the engineered environment that turns a model from a responder into a participant in real work.

The model is only one component. It does not arrive knowing which documents matter, which records are current, which tool is safe, what should persist, or whether a task was actually completed.

The harness makes those decisions explicit. It retrieves and ranks information, assembles a usable context, carries forward selected memory, exposes bounded capabilities, manages the sequence of work, checks the output, and knows when to stop—or when to bring in a person.

That surrounding architecture is the difference between an impressive conversation and an agentic system a company can trust with consequential work.

02The failure map

When production breaks, look outside the model.

A capable model can still fail inside a weak system. Reliability is distributed across the entire harness.

01

Wrong context

The model never receives the record, policy, or prior decision that the task depends on.

Harness fix · retrieval + rankingAt Feniex · Eve checks the exact record before quoting a price or warranty
02

Buried signal

Too much material crowds the working window and important evidence loses priority.

Harness fix · context shapingAt Feniex · a bounded window, short notes, fresh lookups
03

Lost state

The agent forgets task progress, conventions, or decisions and starts reconstructing them.

Harness fix · selective memoryAt Feniex · one conversation record across every door
04

Unsafe action

A tool is available without sufficient limits, permissions, previews, or approvals.

Harness fix · bounded executionAt Feniex · four public actions, and drafting is not sending
05

No proof

The loop ends because the model sounds finished—not because the outcome was verified.

Harness fix · evaluation + stopping rulesAt Feniex · Eve says only what the result proves
03The architecture

Reliability is a pipeline.

Each layer transforms raw capability into usable signal, controlled action, and evidence that the job is done.

01 · Sources

Connect

Documents, code, tickets, systems of record, and live operating data.

02 · Retrieval

Find

Search, filter, rank, and attach metadata to the best evidence.

03 · Context

Shape

Fit instructions, evidence, state, and tool output into working memory.

04 · Model

Reason

Interpret the task, form a plan, choose the next useful step.

05 · Tools

Act

Read, calculate, draft, execute, or interact within explicit boundaries.

06 · Memory

Carry

Persist only the decisions and state that improve future performance.

07 · Evaluation

Verify

Test completion, quality, policy, cost, and whether another loop is needed.

04Retrieval

The first reliability problem is getting the right evidence.

Enterprise knowledge is larger than a prompt. The harness must decide what to retrieve, how to rank it, and what metadata makes it trustworthy.

Search is not one technique.

Lexical retrieval is strong when exact terms matter: a product ID, file name, error code, or contract clause. Dense retrieval is better at meaning: synonyms, paraphrases, and incomplete descriptions. In enterprise work, a hybrid approach often produces the strongest candidate set, followed by re-ranking.

A customer-facing harness adds one rule: look up the exact record before quoting it. In Deyaf’s design, the model is one step of four: listen, look it up, check the rules, reply.

Better evidence creates more useful reasoning before the model itself changes at all.
Retrieval lab · sample queryRelative signal · illustrative
renewal exposure for Orion accounts
Lexical
0.58
Semantic
0.76
Hybrid
0.92
Then re-rank by relevance, recency, authority, permissions, document hierarchy, and task intent.
05Context engineering

Context is finite. Treat it like a budget.

The harness decides what enters the working window, what stays prominent, what is summarized, and what is left out.

A deliberate allocation

An illustrative allocation, not a measurement. There is no universally perfect mix. The useful principle is intentional composition.

Too little context leaves the agent guessing. Too much can bury the reason it was called.

What earns a place

Useful context is not merely related. It is actionable, attributable, and proportionate to the task.

RelevanceDoes it directly change the next decision?
AuthorityIs this the source the business treats as true?
RecencyIs the information current enough to act on?
PermissionMay this identity see and use the record?
StructureCan the model distinguish evidence, instruction, and state?
In Eve’s live answers at Feniex, that budget feeds one step of four, and a finished reply takes about three seconds in website chat and on the phone (Feniex’s own measurements, Sep 17–24, 2026).
06The working system

Memory, tools, and orchestration make intelligence operational.

These layers let an agent carry state, affect the world, and move through work in a controlled sequence. In a customer-facing harness, memory alone splits into four kinds, each with one job.

Layer A

Memory

Preserve what improves future work. Discard what only adds noise or risk.

  • User and project preferences
  • Decisions and task state
  • Reusable procedures
  • Provenance and version history
Layer B

Tools

Give the model capabilities with the narrowest permissions that still make the task useful.

  • Read and search systems
  • Calculate and transform
  • Draft and preview work
  • Execute behind approval gates
Layer C

Orchestration

Decide what happens next, when to retry, how to route, and where a person enters.

  • Plan and execute
  • Route by task type
  • Set cost and time ceilings
  • Stop on proof, risk, or uncertainty
07The operating loop

A reliable agent does not simply continue. It earns the next step.

Agentic retrieval can search, inspect, reformulate, and try again. That improves quality—but also spends time, tokens, and operating budget. The harness makes every loop conditional.

The stop rule matters as much as the start instruction.

Improvement is a loop too, and it needs the same discipline: the learning loop behind Eve measures her before anything new is taught, then measures again.

Policy
+ state
1 · ObserveRead task and evidence
2 · PlanChoose next useful step
3 · ActUse a bounded tool
4 · EvaluateCheck result and risk
5 · DecideStop, retry, or escalate
Architecture discipline Use one agent until the task truly requires coordination. Multiple agents add routing, latency, tokens,
and new failure surfaces.
08Governance

Useful autonomy needs visible boundaries.

The harness is where permission, policy, approval, observability, and accountability become part of the workflow—not a document beside it.

The questions the harness must answer

Good governance is operationally specific. It names who, what, when, and how the system proves compliance. Deyaf publishes its own version as trust and control principles, and its autonomy ladder starts from one line: answering is not acting.

01
Which identity is acting, and what may it see?
02
Which tools are allowed for this task and this moment?
03
What can be drafted, and what can actually be changed?
04
What evidence must appear before a person approves?
05
What gets logged, measured, revoked, or rolled back?
01
IdentityEvery request resolves to a real user or service.
Who
02
PermissionAccess and tools narrow to the current task.
May
03
ProvenanceAnswers and actions retain their source trail.
Why
04
ApprovalConsequential writes stop at a named person.
Gate
05
AuditDecisions, costs, outcomes, and overrides stay visible.
Proof
09In the field

The harness is already answering customers.

Eve is Feniex’s assistant, and each layer in the pipeline above has a counterpart in how she works. Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop.

Illustration: five glowing doorways in mint, periwinkle and amber, each sending a thread of light to one mint point at the center of hairline orbital rings, on a navy field.
Plate · One assistant, every doorIllustration · mood only

Eve is Feniex’s warm, composed public face. She helps people choose suitable equipment, make a buying decision, or get support for equipment they already own. She answers first, gives one useful next step, then stops. On the phone she is the overflow line: the team’s phones ring first, and when no one is free or the office is closed, Eve answers instead of voicemail.

Her answers are assembled, not improvised. Before any exact product, compatibility, warranty, price or software claim, she checks the applicable record: a versioned reference layer of catalog, manual, software, fitment and policy records, then a taught library of reviewed lessons found by meaning. A convincing product name is not evidence.

When she doesn’t know, the harness takes over, not the model’s imagination. The unanswered question becomes a durable item for the team, and a support ticket is filed when contact details are available. She says it is with the support team only once the ticket is confirmed as accepted. There are no transfers, by design: on a call she takes a name and number and hands the matter on. That is what happens when she doesn’t know.

What she remembers is split by job: this conversation, one conversation record the team can read across web, email, text and phone, short customer notes, and her library. Notes are never current facts; orders, invoices, shipments and balances are always looked up fresh. Raw conversations never become public knowledge on their own. Your team answers once, the lesson is reviewed, she is measured on a frozen test set, only explicitly approved lessons are published, and then she is measured again.

The five controls, at Feniex

The governance stack from chapter 08, as it reads in Eve’s harness.

WhoRecognizes a returning customer’s account context, limited to that one account.
MayFour public actions: look something up, search the taught library, take a message, record an unanswered question.
WhyThe exact reference layer, checked before any exact claim.
GateDrafting is not sending. She claims no authority to approve returns, change accounts or place orders.
ProofShe describes only what the result proves: saved, drafted, sent, uncertain, acknowledged or resolved.
Autonomy is set per workflow and per channel, on a published three-level ladder: answer only, prepare for approval, run an approved task.

Measured, not promised

Graded Sep 23, 2026 on 100 fixed test questions. Each bar counts questions out of 100.

86% handled well, with a 10% material-error rate and 0 critical errors. The trend ran 88% (Sep 17) → 90% (Sep 21) → 86% (Sep 23), and past trouble spots scored 40%, a group kept in the test on purpose.

As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional.

Eve’s report card, weak spots included
10At the door

Where the system meets the customer.

Every layer above stays out of sight. What a customer actually meets is a small box on a web page and a phone line that rings. The same identity, library and rules stand behind both; only the delivery changes at each door.

Door A · Website chat

The chat box

The smallest, most visible door. The reply streams in as it is written, the page and the product in view ride along, and an answer can carry a structured product or resource tile. Its history is bounded to this one conversation.

  • Answers first, gives one useful next step, then stops
  • Asks for one missing detail at a time
  • Checks the exact record before any product, warranty, price or software claim
  • Never pretends to be a person
Door B · Phone

The phone line

Overflow and after hours. The team’s phones ring first; when no one is free or the office is closed, Eve answers instead of voicemail: Hello, this is Eve, Feniex’s intelligence. How can I help you?

  • Never silent: “One moment” at about 1.5 s, “Let me check on that” as a lookup starts, “Still checking” about every 4 s
  • Callers in Spanish, French or Portuguese hear that language’s own voice
  • Every call transcribed, a recorded fallback if she is down, 10 minutes at most
  • Never reads a web address aloud
Routing, both ways

No transfers, by design. Routing in is overflow: the team first, then Eve. Routing out is a hand-off of details, never a transfer: she answers from her library, or takes the caller’s name and number and hands the matter to the team. In chat, an unanswered question becomes a durable item, and she says “with the support team” only once the ticket is confirmed accepted. How Eve works the overflow line

11Enterprise scale

The architecture should scale in control—not just volume.

Control describes the discipline of the system, not the size of the customer. A smaller company can need the same provenance, permissions, and human control as a global one.

Small enterprise

One doorway into scattered knowledge.

Start with a focused outcome and a few trusted systems. Preserve source, ownership, and approval from the first day.

Start with a front desk that knows↗
Global enterprise

Wide intelligence. Narrow authority.

Operate across complex systems while keeping access local, writes bounded, and every consequential action accountable.

What the assistant can and cannot do↗
12The field checklist

Harness engineering asks better systems questions.

Before debating whether the model is smart enough, inspect the operating environment it has been given. Then grade it the way Eve is graded: on a fixed test, hard cases kept in.

01

Information

What evidence does this task require, who owns it, and how will retrieval quality be measured?

02

Context

What must remain visible, what can be summarized, and what should never enter the working window?

03

Memory

Which decisions improve future work, how long should they persist, and who can correct them?

04

Capability

Which tools are necessary, what is the minimum permission, and where must the system preview before acting?

05

Orchestration

When should the agent retry, reformulate, route, ask a person, or stop?

06

Evaluation

What observable evidence proves the outcome is correct, complete, safe, and worth the operating cost?

07

Governance

Which identities, policies, approvals, and audit events must be inside the workflow?

08

Operations

How will quality, latency, cost, permissions, and failures be monitored over time?

Research note

This editorial interpretation draws on the ODSC article “What is an Agent Harness? The Architecture Behind Reliable Agentic AI”. The architecture, diagrams, commentary, and Deyaf framing on this page are original to Site Number Three.

Further reading

These sources describe the wider field. None of them reviewed or endorses Deyaf, and none describes how Eve is built.

Build the system, not just the demo.

Deyaf packages the harness behind Eve so your business can have an assistant of its own. Start in the builder: about ten minutes to a knowledge preview and a saved setup, no account needed.

Build your assistant