Identity
Defines the agent’s role, priorities, communication style, boundaries, and stop conditions.
That's the Agentic Harness.
A model can reason. A production system must also remember, choose tools, recover from failure, show its work, and know when to stop. The Agentic Harness is the engineered runtime that makes those behaviors repeatable.
Editorial starting point. This original field guide builds on the production architecture described by Mrunmayee Rane in “The Agentic Harness: How to Build AI Agents in Production”. It extends that technical frame into a business-readable system map.
What the harness surrounds
The source article makes a useful architectural move: responsibilities that are often crammed into a prompt are externalized into owned layers. That makes the intelligence easier to control, inspect, replace, and improve. Deyaf’s definition says it in one sentence: “An Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.” How an Agentic Harness works
Defines the agent’s role, priorities, communication style, boundaries, and stop conditions.
States how work begins, what “done” means, how progress is checked, and how risk is reported.
Routes fresh, relevant working information to the model without treating the context window as storage.
The model is one step of fourPreserves durable facts and decisions outside the conversation so long-running work can continue.
Four kinds of memory, one job eachThe model interprets the task, chooses a next step, and adapts when a rigid path cannot.
Tools expose controlled actions. Skills encode reusable procedures. Policies determine permitted use.
Answering is not actingCheckpoints, retry rules, progress artifacts, and escalation paths prevent silent failure and lost work.
Traces reveal the trajectory. Evals grade both outcome and process. Governance gives permissions, changes, and incidents a clear owner.
The production loop
Production quality comes from the execution loop living outside the model. The harness supplies continuity, policies, and evidence at each transition.
Capture the request, outcome, user identity, risk, and scope.
Retrieve the right instructions, memory, sources, and procedures.
Route simple work to code; flexible work to an agent loop.
Execute within permissions, budgets, retry limits, and approval rules.
Inspect results, test the outcome, and recover or stop when needed.
Write durable state, record the trajectory, and feed evaluation.
The model reasons · the harness routes, constrains, remembers, verifies, and records
Choose the simplest shape
A healthy architecture matches agency to uncertainty. The more freedom a system has to choose its path, the more policy, observation, evaluation, and human control it needs around that freedom.
One bounded input produces one bounded output. Fast, economical, and easy to test.
The system fetches relevant knowledge before the model answers, with sources available for inspection.
Code owns the sequence. AI can help inside a step, but it does not invent the process.
The system researches or prepares; a named person reviews consequential action before execution.
The model chooses steps and tools across a bounded loop, adapting as evidence changes.
Specialized agents work in parallel or review one another under explicit handoff contracts.
An industry map of system shapes, not a product ladder. Deyaf’s live product publishes three levels of autonomy instead: answer only, prepare for approval, and run an approved task. Each task and each channel moves up separately, and there is no single switch that lets an assistant do everything. See the three levels.
Production is a reliability problem
A polished answer can hide a bad trajectory: the wrong tool, stale context, needless retries, unsafe edits, or an unverified result. The harness makes that invisible work inspectable.
What to measure, defined, not scored. For one dated set of real numbers, see Eve at Feniex.
Capture context, tool calls, outputs, state transitions, retries, latency, cost, and feedback.
Grade how the work was done—not only whether the final response looks correct.
Define retry limits, safe alternatives, escalation points, and conditions that require a stop.
Give identities, permissions, memories, skills, releases, and incidents a clear owner.
Keep consequential writes visible, explainable, and gated by the right person.
The harness, already answering
Eve is Feniex’s assistant, and the clearest way to watch this guide’s map at work. Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. Meet Eve, Feniex’s assistant
Illustrative example: not a live Eve transcript. One question walking down the layers of this guide’s map.
Feniex’s warm, composed public face (Feniex is said like “Phoenix”). She helps people choose suitable equipment, make a buying decision, or get support for equipment they own. She is software with a named personality, not a person, and approved Feniex operators manage her rules and knowledge.
One Eve, every door: website chat, website voice, phone, text and email, in English, Spanish, French, Portuguese and Arabic. On the phone she is the overflow and after-hours line. The team’s phones ring first; when no one is free or the office is closed, Eve answers instead of voicemail.
One assistant, every way inOne library behind every answer: versioned catalog, manual, software, fitment and policy records, plus lessons the team has reviewed. She checks the applicable record before any exact product, warranty, price or compatibility claim, because a convincing product name is not evidence.
Inside one answer, step by stepFour public actions: look something up, search the taught library, take a message, and record an unanswered question. She describes only what the result proves (saving is not delivery, delivery is not resolution) and claims no authority to approve returns, change accounts or place orders.
What she may do, level by levelUnknown stays unknown. The question becomes a durable unanswered item and, when contact details are available, a support ticket. She says it is with the support team only once the ticket is confirmed. No transfers, by design: on a call she takes the caller’s name and number and hands the matter to the team.
What happens when she doesn’t knowFour kinds of memory, each with one job: this conversation, one conversation record the team can read, short customer notes, and her reviewed library. Orders, invoices and balances are always looked up fresh, never recalled from notes. Raw conversations never become public knowledge on their own; only lessons a person approves are taught.
As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional. These are Feniex’s numbers, not a forecast: a new assistant starts fresh with its own knowledge.
Where customers meet the harness
Everything in this guide runs out of sight. Customers meet it in two places: a small chat box on a web page, and a phone line that rings when the team can’t pick up. The same identity, library and rules stand behind both; only the delivery changes.
Calls since Sep 17, 2026, her first week on the line. As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional. Eve’s report card, speed included
The smallest, most visible door. Replies stream in as they are written, the page and product in view ride along, an answer can carry a structured product or resource tile, and history stays inside this conversation.
She answers first, gives one useful next step, then stops, and asks for one missing detail at a time. The exact record comes before any product, warranty, price or software claim. She never pretends to be a person.
Mind Your Manners: chat-box etiquetteThe team’s phones ring first. When no one is free or the office is closed, Eve answers instead of voicemail: Hello, this is Eve, Feniex’s intelligence. How can I help you?
Callers in Spanish, French or Portuguese hear that language’s own voice.
“One moment” at about 1.5 s, “Let me check on that” as a lookup starts, “Still checking” about every 4 s. Every call is transcribed, a recorded fallback plays if she is down, and she never reads a web address aloud.
The Elements of a Phone Assistant, as a periodic tableNo transfers, by design. In: overflow, the team first, then Eve. Out: she answers from her library, or takes the caller’s name and number and hands the matter to the team. In chat, she says “with the support team” only once the ticket is confirmed accepted.
From a working assistant to yours
Deyaf is built from Eve: the harness that runs Eve at Feniex, packaged so another business can have an assistant of its own. The current release is an early-access setup and preview experience: a four-step builder (your business, what it knows, how it helps, try it) with an exact-source knowledge preview (browser-local, not live AI). Live doors switch on one at a time, each one deliberately.
See the harness behind EveChoose a bounded, measurable business result with a responsible owner and an honest baseline.
Use a model call, retrieval, or fixed workflow when it is enough. Add an agent only when adaptation pays.
Decide what the system may know, where truth lives, who can access it, and what must never enter context.
Give every tool an allowed scope, preconditions, failure behavior, retry limit, logging rule, and approval boundary.
Trace decisions and state changes. Evaluate repeat runs for success, consistency, cost, safety, and recovery. Measure on a fixed test set before anything new is taught. Measured before anything is taught
Promote capability only after the system repeatedly earns trust under realistic conditions. Switch on one door, one task at a time. Switch on each door deliberately
Further reading
The production architecture this field guide starts from.
Workflows versus agents, and why the simplest shape that works usually wins.
Agent equals model plus harness, with each harness piece derived from that split.
State and recovery in practice: feature lists, progress files, and clean handovers between sessions.
Capability versus regression evals, and why consistency across repeat runs matters.
Tracing agent runs as spans for the agent, the model call, and each tool.
Guides and sensors: steering an agent before it acts, and checking its work after.
Independent industry sources. None of them endorses Deyaf or describes Eve’s stack.
Deyaf’s builder walks through four steps and ends in an exact-source knowledge preview (browser-local, not live AI). Try it in about ten minutes, no account needed.
Build your assistant Or start with one good job