Systems essay 04 10 min read · Updated Sep 2026

The work lives around the model.

That's the Agentic Harness.

A model can reason. A production system must also remember, choose tools, recover from failure, show its work, and know when to stop. The Agentic Harness is the engineered runtime that makes those behaviors repeatable.

Editorial starting point. This original field guide builds on the production architecture described by Mrunmayee Rane in “The Agentic Harness: How to Build AI Agents in Production”. It extends that technical frame into a business-readable system map.

01The structure

What the harness surrounds

The model is one layer. The harness is seven more.

The source article makes a useful architectural move: responsibilities that are often crammed into a prompt are externalized into owned layers. That makes the intelligence easier to control, inspect, replace, and improve. Deyaf’s definition says it in one sentence: “An Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.” How an Agentic Harness works 

L0

Identity

Defines the agent’s role, priorities, communication style, boundaries, and stop conditions.

L1

Task contract

States how work begins, what “done” means, how progress is checked, and how risk is reported.

L4

Reasoning

The model interprets the task, chooses a next step, and adapts when a rigid path cannot.

L5

Tools + skills

Tools expose controlled actions. Skills encode reusable procedures. Policies determine permitted use.

Answering is not acting 
L6

State + recovery

Checkpoints, retry rules, progress artifacts, and escalation paths prevent silent failure and lost work.

02How it runs

The production loop

A useful answer is not enough. The path must hold up.

Production quality comes from the execution loop living outside the model. The harness supplies continuity, policies, and evidence at each transition.

01 / receive

Frame the task

Capture the request, outcome, user identity, risk, and scope.

02 / load

Assemble context

Retrieve the right instructions, memory, sources, and procedures.

03 / decide

Choose the path

Route simple work to code; flexible work to an agent loop.

04 / act

Use a tool

Execute within permissions, budgets, retry limits, and approval rules.

05 / prove

Verify progress

Inspect results, test the outcome, and recover or stop when needed.

06 / learn

Persist + trace

Write durable state, record the trajectory, and feed evaluation.

The model reasons · the harness routes, constrains, remembers, verifies, and records

03System types

Choose the simplest shape

Not every AI feature should become an agent.

A healthy architecture matches agency to uncertainty. The more freedom a system has to choose its path, the more policy, observation, evaluation, and human control it needs around that freedom.

Type 01Agency / none

Single model call

One bounded input produces one bounded output. Fast, economical, and easy to test.

Best fitClassification, extraction, rewriting, structured summaries
Type 02Agency / low

Retrieval system

The system fetches relevant knowledge before the model answers, with sources available for inspection.

Best fitPolicy Q&A, knowledge search, grounded support
Type 03Agency / fixed

Deterministic workflow

Code owns the sequence. AI can help inside a step, but it does not invent the process.

Best fitRepeatable operations, compliance-heavy flows, known decisions
Type 04Agency / gated

Human-in-the-loop

The system researches or prepares; a named person reviews consequential action before execution.

Best fitCustomer messages, financial changes, hiring, publishing
Type 05Agency / adaptive

Single agent

The model chooses steps and tools across a bounded loop, adapting as evidence changes.

Best fitInvestigation, coding, research, exception handling
Type 06Agency / coordinated

Multi-agent system

Specialized agents work in parallel or review one another under explicit handoff contracts.

Best fitBroad research, independent verification, large system analysis

An industry map of system shapes, not a product ladder. Deyaf’s live product publishes three levels of autonomy instead: answer only, prepare for approval, and run an approved task. Each task and each channel moves up separately, and there is no single switch that lets an assistant do everything. See the three levels.

04Reliability

Production is a reliability problem

Measure the journey, not just the destination.

A polished answer can hide a bad trajectory: the wrong tool, stale context, needless retries, unsafe edits, or an unverified result. The harness makes that invisible work inspectable.

PassTask success and consistency
TraceTool and state trajectory
CostTokens, tools, time per success
RecoverFailure and escalation behavior

What to measure, defined, not scored. For one dated set of real numbers, see Eve at Feniex.

01

Observability

Capture context, tool calls, outputs, state transitions, retries, latency, cost, and feedback.

02

Trajectory evaluation

Grade how the work was done—not only whether the final response looks correct.

03

Failure policy

Define retry limits, safe alternatives, escalation points, and conditions that require a stop.

04

Governance

Give identities, permissions, memories, skills, releases, and incidents a clear owner.

05

Human control

Keep consequential writes visible, explainable, and gated by the right person.

05In production

The harness, already answering

The layers are not theory. They already answer.

Eve is Feniex’s assistant, and the clearest way to watch this guide’s map at work. Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. Meet Eve, Feniex’s assistant 

Flat poster illustration of a small desk with a lime lamp at the center of an open, roofless room, with thin orange lines running from the desk to every doorway in its walls.
One desk behind every door. Illustration for mood only.

Illustrative example: not a live Eve transcript. One question walking down the layers of this guide’s map.

01 / Who

Who she is

Feniex’s warm, composed public face (Feniex is said like “Phoenix”). She helps people choose suitable equipment, make a buying decision, or get support for equipment they own. She is software with a named personality, not a person, and approved Feniex operators manage her rules and knowledge.

02 / Doors

Where customers reach her

One Eve, every door: website chat, website voice, phone, text and email, in English, Spanish, French, Portuguese and Arabic. On the phone she is the overflow and after-hours line. The team’s phones ring first; when no one is free or the office is closed, Eve answers instead of voicemail.

One assistant, every way in 
03 / Knows

What she knows

One library behind every answer: versioned catalog, manual, software, fitment and policy records, plus lessons the team has reviewed. She checks the applicable record before any exact product, warranty, price or compatibility claim, because a convincing product name is not evidence.

Inside one answer, step by step 
04 / May do

What she may do

Four public actions: look something up, search the taught library, take a message, and record an unanswered question. She describes only what the result proves (saving is not delivery, delivery is not resolution) and claims no authority to approve returns, change accounts or place orders.

What she may do, level by level 
05 / Unknown

When she doesn’t know

Unknown stays unknown. The question becomes a durable unanswered item and, when contact details are available, a support ticket. She says it is with the support team only once the ticket is confirmed. No transfers, by design: on a call she takes the caller’s name and number and hands the matter to the team.

What happens when she doesn’t know 
06 / Memory

What she remembers

Four kinds of memory, each with one job: this conversation, one conversation record the team can read, short customer notes, and her reviewed library. Orders, invoices and balances are always looked up fresh, never recalled from notes. Raw conversations never become public knowledge on their own; only lessons a person approves are taught.

Report card86%handled well on 100 fixed test questions, graded Sep 23, 2026, with a 10% material-error rate and 0 critical errors.
Eve’s full report card 
Phone102calls answered in her first week on the line: 86 while the team was busy, 16 after the office closed. Sep 17–24, 2026.
Eve on the phone 
Hand-offs0 lostSep 16–23, 2026. Every one sent reached the team.

As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional. These are Feniex’s numbers, not a forecast: a new assistant starts fresh with its own knowledge.

06At the door

Where customers meet the harness

Two doors customers see. One harness behind them.

Everything in this guide runs out of sight. Customers meet it in two places: a small chat box on a web page, and a phone line that rings when the team can’t pick up. The same identity, library and rules stand behind both; only the delivery changes.

~3 sTo a finished reply, in chat and on the phone
93%Of callers stayed and talked
60Calls handed straight to the team, details taken
10 minThe longest a call runs

Calls since Sep 17, 2026, her first week on the line. As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional. Eve’s report card, speed included

02

Chat manners

She answers first, gives one useful next step, then stops, and asks for one missing detail at a time. The exact record comes before any product, warranty, price or software claim. She never pretends to be a person.

Mind Your Manners: chat-box etiquette 
03

The phone line

The team’s phones ring first. When no one is free or the office is closed, Eve answers instead of voicemail: Hello, this is Eve, Feniex’s intelligence. How can I help you? Callers in Spanish, French or Portuguese hear that language’s own voice.

Nobody Has to Hold: automatic call taking 
04

Never silent

“One moment” at about 1.5 s, “Let me check on that” as a lookup starts, “Still checking” about every 4 s. Every call is transcribed, a recorded fallback plays if she is down, and she never reads a web address aloud.

The Elements of a Phone Assistant, as a periodic table 
DEYAF.

From a working assistant to yours

Built from Eve. Packaged for your business.

Deyaf is built from Eve: the harness that runs Eve at Feniex, packaged so another business can have an assistant of its own. The current release is an early-access setup and preview experience: a four-step builder (your business, what it knows, how it helps, try it) with an exact-source knowledge preview (browser-local, not live AI). Live doors switch on one at a time, each one deliberately.

See the harness behind Eve
Eve at FeniexAnswering Feniex’s customers in production today.
Watch Eve answer on five doors 
Five doorsWebsite chat, voice, phone, text and email, with one library behind them.
Five doors, one library 
Five languagesEnglish, Spanish, French, Portuguese and Arabic, answered from one English library.
How Eve handles five languages 

Name one outcome

Choose a bounded, measurable business result with a responsible owner and an honest baseline.

Choose the system type

Use a model call, retrieval, or fixed workflow when it is enough. Add an agent only when adaptation pays.

Map context + authority

Decide what the system may know, where truth lives, who can access it, and what must never enter context.

Constrain tools

Give every tool an allowed scope, preconditions, failure behavior, retry limit, logging rule, and approval boundary.

Instrument the loop

Trace decisions and state changes. Evaluate repeat runs for success, consistency, cost, safety, and recovery. Measure on a fixed test set before anything new is taught. Measured before anything is taught 

Widen on evidence

Promote capability only after the system repeatedly earns trust under realistic conditions. Switch on one door, one task at a time. Switch on each door deliberately 

08Questions

Questions worth asking.

Is an Agentic Harness a new model?
No. The model supplies reasoning. The harness is the runtime around it: identity, context, memory, tools, state, policies, verification, traces, evaluation, and governance.
Does every AI system need an agent?
No. Many useful systems should remain a single model call, a retrieval flow, or a deterministic workflow. Agency is valuable when the path cannot be fully known in advance.
What is the difference between memory and a skill?
Memory stores durable facts—what the system knows. A skill stores a reusable procedure—how the system performs a kind of work.
Why is a human gate still important?
Reading and preparing create value with limited operational risk. Consequential writes can affect customers, money, people, or records. A named approval boundary keeps responsibility visible.
Where does Deyaf fit?
Deyaf packages the harness that already runs Eve at Feniex, so a business can set up an assistant of its own: its knowledge, its rules, and the doors it chooses to open, one at a time. The starting plan requires human review before external messages, record changes, or commitments, and every task begins at the first of three levels: answering is not acting.

Further reading

  1. The Agentic Harness: How to Build AI Agents in ProductionMrunmayee Rane · dev.to

    The production architecture this field guide starts from.

  2. Building Effective AI AgentsAnthropic · Dec 2024

    Workflows versus agents, and why the simplest shape that works usually wins.

  3. The Anatomy of an Agent HarnessLangChain · Mar 10, 2026

    Agent equals model plus harness, with each harness piece derived from that split.

  4. Effective harnesses for long-running agentsAnthropic · Nov 2025

    State and recovery in practice: feature lists, progress files, and clean handovers between sessions.

  5. Demystifying evals for AI agentsAnthropic · Jan 2026

    Capability versus regression evals, and why consistency across repeat runs matters.

  6. GenAI observabilityOpenTelemetry · 2026

    Tracing agent runs as spans for the agent, the model call, and each tool.

  7. Harness engineering for coding agent usersBirgitta Böckeler · martinfowler.com · Apr 2, 2026

    Guides and sensors: steering an agent before it acts, and checking its work after.

Independent industry sources. None of them endorses Deyaf or describes Eve’s stack.

Build the system around the intelligence.

Deyaf’s builder walks through four steps and ends in an exact-source knowledge preview (browser-local, not live AI). Try it in about ten minutes, no account needed.

Build your assistant Or start with one good job