Field atlas 08 · 12 min readSystems / reliability / control

The model is one node. Reliability is the territory.

Agentic Harness

A language model can reason. An Agentic Harness gives that reasoning somewhere safe to work: instructions, tools, memory, execution, observation, validation, permissions, and people.

This original visual explainer builds on Databricks’ June 2026 guide to Agentic Harnesses, maps the operating decisions organizations face beyond the model, then follows one working harness, Eve at Feniex, across the same eight territories.

How an Agentic Harness works, end to end
Interactive system mapSelect a territory
ModelReasons
and decides
01Territory
Instructions set the standing brief.

The harness supplies the role, objective, boundaries, and durable rules before a task begins.

In EveHer identity and rules load first: answer, offer one useful next step, then stop.

The useful equation

Do not ask the model to be the whole system.

Databricks separates three ideas teams often blur together. The distinction changes what you build, what you test, and where you look when production work breaks.

01 / Model
Reasons.

Interprets context, predicts, plans, and produces the next decision or output.

02 / Harness
Operates.

Supplies context and tools, executes actions, preserves state, checks work, enforces rules, and records what happened.

03 / Agent
Works.

The complete system a person experiences: reasoning joined to controlled execution.

Operational consequence: model selection is only one design decision. The surrounding harness determines what the system may see, do, remember, verify, and escalate. Deyaf’s definition says it in customer terms: “An Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.”

See the four steps of one answer

The control loop

Reliable action is a cycle, not a leap.

The harness carries decisions into an environment, captures the result, and returns evidence to the model. Validation and human oversight keep repetition from becoming runaway autonomy.

The loop has two authors.

The model chooses what to try. The harness defines how the attempt reaches the world and what comes back. A human or policy gate can decide whether consequential work crosses the final boundary.

  1. Reasoning stays revisable. New observations can change the plan.
  2. Execution stays bounded. Tools and sandboxes limit where an attempt can land.
  3. Evidence stays visible. Tests, traces, and results make the next step inspectable.
  4. Authority stays explicit. High-impact actions can wait for a named approver.

The interleaving of reasoning and action was formalized in the 2022 ReAct paper. The approval and validation layer shown here is an editorial extension for operational use. Deyaf publishes its version as a three-level ladder, from answer only to one approved task at a time: answering is not acting.

Illustrative example

One run, traced.

One hypothetical run of an order-support agent, not a customer transcript and not an Eve log. Each pass around the loop leaves a line someone can inspect later.

StepWhat happenedRouteWhat the trace records
01 / ModelReason
A customer asks to change the delivery address on an open order. The model reads the request, the order context and the rule that address changes need confirmation, and plans a lookup first.
Instructions + sources

Which rule version and which retrieved records were in context.

02 / HarnessAct
The harness calls the order-lookup tool with a read-only credential scoped to this one customer.
Tool call + scope

The tool name, its arguments, and the permission level it ran under.

03 / HarnessObserve
The lookup comes back: order found, not yet shipped. That result returns to the model as evidence.
Result + timing

What came back, how long it took, and whether anything failed.

04 / SystemValidate
Changing an address is a write with consequences, so the run pauses and routes a prepared change to a named approver.
Rule + approver

Which policy fired, and who now holds the decision.

05 / PersonApprove
The approver says yes. Only then does the harness apply the change and confirm it back to the customer.
Authority + outcome

Who authorized the write, when, and exactly what changed.

Eight production territories

The harness is not one wrapper. It is a stack of responsibilities.

Databricks identifies eight building blocks. Read together, they form a route from intention to action—and back to evidence.

01
ISystem instructions
Sets the standing brief. Role, objective, boundaries, and rules enter before the task.
02
TTools & execution
Turns decisions into calls. Search, APIs, databases, code, and business applications become usable capabilities.
03
XSandbox
Contains experiments. Generated code and uncertain actions run inside an isolated environment.
04
FFilesystem & storage
Gives work a durable place. Files, plans, and intermediate artifacts can persist beyond one message.
05
MMemory & context
Curates what remains active. The harness retrieves, summarizes, and carries forward relevant history.
06
VFeedback & verification
Checks before claiming success. Tests, inspection, and review loops reveal incomplete or incorrect work.
07
GGuardrails & humans
Defines authority. Policy blocks, approvals, and human review protect consequential boundaries.
08
OObservability & logs
Makes behavior inspectable. Traces show what happened, why, where it failed, and under whose authority.

Editorial interpretation: the eight layers are best treated as one control surface. A memory improvement can create a privacy risk; a new tool can create a permission problem; more autonomy raises the value of verification and logs.

Risk-to-guardrail matrix

Every failure mode is a design request.

The source names common production failures. The useful response is not “use a smarter model.” It is to place a concrete control at the point where the system can drift.

Failure modeWhat the operator seesRouteHarness control
Context rot
Long tasks become noisy; older material crowds out what matters now.
Compaction + retrieval

Summarize stale history, retain decisions, and retrieve only relevant evidence.

Irrelevant retrieval
Search returns plausible but off-target material, and the model treats it as evidence.
Scoped sources + ranking

Filter retrieval to the task, rank for relevance, and let “not found” stay an honest answer.

Tool overload
The model spends effort choosing among too many possible capabilities.
Scoped discovery

Expose the smallest capable toolset for the current task and authority.

Brittle wiring
A small schema or description change produces silent tool misuse.
Contracts + tests

Version tool interfaces, validate arguments, and run integration checks.

Weak verification
The agent declares completion before the outcome is actually correct.
Outcome gates

Require tests, inspection, or an independent review step before completion.

Missing guardrails
An irreversible message, deletion, or purchase crosses the boundary unchecked.
Permission + approval

Enforce least privilege and route high-impact actions to a named person.

Latency
Every extra call, retry, and check adds waiting until a person gives up on the answer.
Step budgets

Time-limit each step, run independent work in parallel, and keep the common path short.

Control recommendations are an editorial synthesis, not measured claims. NIST’s voluntary AI Risk Management Framework similarly treats trustworthiness as a consideration across system design, development, use, and evaluation.

Atlas in the field

The same map, already at work: Eve at Feniex.

Every territory above has a working counterpart in Eve, Feniex’s assistant. Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop.

Meet Eve, Feniex’s assistant
Field plate · Eve at FeniexEvery door
Mood illustration: a blank survey plate on graph paper, with ink contour lines rippling out from a single red point and thin dashed red routes running to doorways cut into the frame.
Mood illustration, not a diagram: one red point, many ways in.

Who she is, and where customers reach her.

Eve is Feniex’s warm, composed public face: a knowledgeable customer desk that helps people choose suitable equipment, make a buying decision, or get support for equipment they own. She is software with a named personality, not a person, and approved Feniex operators manage her personality, rules and knowledge.

One Eve, every door: website chat, website voice, phone, text and email, in five languages answered from one English library. On the phone she is the overflow and after-hours line. The team’s phones ring first; when no one is free, she answers instead of voicemail. No transfers, by design: she answers from her library or takes a name and number for the team.

One assistant, every way in
01
ISystem instructions
Identity and rules. Her name, manner and rules load before every conversation: answer first, ask for one missing detail at a time, never add a sales pitch to support. How she talks, on Meet Eve
02
TTools & execution
Four public actions. Look something up, search the taught library, take a message, record an unanswered question. The tool list and server-side queries enforce that scope. What she may do, one level at a time
03
XSandbox
Walled channels. Unattended replies are read-only, and a text or email can never start extra work. She claims no authority to approve returns, change accounts or place orders. What the assistant can and cannot do
04
FFilesystem & storage
The exact reference layer. Versioned catalog, manual, software, fitment and policy records with verified resource links. She checks the record before any exact product, compatibility, warranty, price or software claim. The model is one step of four
05
MMemory & context
Four kinds of memory. This conversation, the conversation record, short customer notes and her library. Notes are never current facts: orders, invoices and balances are always looked up fresh. Four kinds of memory, each with one job
06
VFeedback & verification
Measured, then taught. Tickets and calls fold into a redacted question book. She is measured on a frozen test set, only approved lessons are published, then she is measured again: a supervised loop, not training on every raw conversation. The loop that makes her better
07
GGuardrails & humans
A human gate and a real hand-off. Consequential steps wait for a person’s yes. What she can’t answer becomes a durable unanswered item, and she says “with the support team” only once the ticket is accepted. What happens when she doesn’t know
08
OObservability & logs
A record and a report card. One operator-readable conversation record across web, email, text and phone, never searchable by visitors, plus a report card graded on 100 fixed test questions. The dated report card
Report card
86%handled well on 100 fixed test questions, graded Sep 23, 2026
Always beside it
10%material errors on the same test, with 0 critical errors
Phone line
102calls answered in her first week on the line, from Sep 17, 2026
Hand-offs
0 lostevery hand-off sent reached the team, Sep 16–23, 2026

As of September 24, 2026 · Feniex’s internal Eve 3.0 report · machine-judged and provisional · Feniex’s numbers, not a forecastEve’s report card, weak spots included

New territory

Where the map meets customers: the chat door and the phone door.

The eight territories sit behind every answer; customers only see the doors. Two are surveyed here, with the same identity, library and rules as website voice, text and email.

D1
CWebsite chat
The smallest door, the most visible. The reply streams in, page and product context ride along, and product or resource tiles can arrive with it. History is bounded to this one conversation. Inside the box in the corner, as a comicFrom scripted bots to chat that answers from sources
D2
MChat manners
Answer, one next step, then stop. Usually two or three sentences, one missing detail asked at a time. She checks the exact record before any product, compatibility, warranty, price or software claim, says only what the result proves, and never pretends to be a person. Chat-box etiquette, from “say it’s AI” to an honest path to a person
D3
PPhone: overflow & after hours
The team first, then Eve. When no one is free or the office is closed, she answers instead of voicemail: “Hello, this is Eve, Feniex’s intelligence. How can I help you?” Never silent: “One moment” at about 1.5 seconds, “Let me check on that” as a lookup starts, “Still checking” about every four seconds. Nobody has to hold: a phone line’s 24 hoursEve on the phone
D4
RPhone: routing
No transfers, by design. She answers from her library, or takes a name and number and hands the matter to the team. Every call is transcribed, a recorded fallback plays if she is down, and calls run ten minutes at most. Spanish, French and Portuguese callers hear their own language’s voice. Press 1 is over: why routing ends in a hand-offThe elements of a phone assistant, as a periodic table

Surveyed from Eve’s published rules · same rules at every doorOne assistant, every way in

One architecture, different emphasis

Scale changes the control problem—not the need for control.

Select an organization size to see an illustrative starting posture. These are planning examples, not customer results or universal prescriptions.

Illustrative example · Small organization

Narrow map. Visible owner.

Start with one valuable workflow, a compact tool surface, and an approval boundary every operator can explain.

Prioritize clarity before breadth.

  • One outcome-shaped workflow with a named business owner
  • Only the tools and data needed for that outcome
  • Human approval before external or irreversible action
  • Simple traces, a stop control, and a tested recovery path
A front desk that knows: try the builder
Illustrative example · Mid-size organization

Shared layer. Clear handoffs.

As teams and systems multiply, reuse the harness infrastructure while keeping workflow ownership and permissions legible.

Standardize the seams between teams.

  • Reusable instructions and tools for repeated operating patterns
  • Central identity with role-aware access to connected systems
  • Shared evaluation and observability across departments
  • Approval routes that follow existing business ownership
Start with one good job, then grow
Illustrative example · Large organization

Many agents. One control plane.

At enterprise scale, shared governance, model flexibility, isolation, evaluation, and auditability become infrastructure concerns.

Treat the harness as governed infrastructure.

  • Central policy for data, tools, models, and execution environments
  • Least privilege, workload isolation, and formal change control
  • Continuous evaluation and end-to-end traces at production volume
  • Auditable human authority for regulated or high-impact actions
Deyaf’s trust and control principles

From atlas to operating plan

Build outward from the work, not inward from the model.

A practical first path connects an outcome to the minimum system around it, then widens only after evidence and ownership are in place.

01 / Outcome

Name the work.

Choose one result, the people it serves, the evidence of success, and the conditions that should stop the system.

02 / Territory

Bind the minimum.

Map only the instructions, context, tools, storage, and environment required for that result.

03 / Authority

Set the boundary.

Define who owns each connection, what the agent may prepare, and what needs explicit human approval.

04 / Evidence

Watch and widen.

Trace runs, test outcomes, review failures, and expand the map only when the control model holds.

Field questions

What people ask after reading the map.

Short answers to the questions that come up once the eight territories are on the table. Editorial answers, not source claims.

Q1

Is a harness the same thing as an agent framework?

Not quite. A framework is a kit for building agents. The harness is the running system around one model doing one job: its instructions, tools, memory, sandbox, checks, permissions and records. You can assemble it with a framework, by hand, or from a packaged product; the eight territories are the same either way.
Q2

If models keep improving, does the harness matter less?

It matters more. A better model raises the ceiling on what one step can do. It does not decide what the system may see, touch, remember or claim, or who approves a consequential action. Those stay harness decisions, and they carry more weight as the agent is allowed to do more.
Q3

What should a small team build first?

One job, one owner. Pick a single workflow, the fewest tools that finish it, read-only access wherever possible, and a person’s approval before anything external or irreversible. Widen the map only after traces and tests show the first loop holds. Start from a support desk template
Q4

How do you know a harness change helped?

Measure before and after on the same test. Keep a fixed set of questions, grade every change against it, and print the error rate beside the success rate. Eve’s report card above follows that rule: its success figure never appears without its material-error figure.
Q5

What should an assistant do when it does not know?

Say so, then hand it on. A good harness turns “I don’t know” into a hand-off a person can act on. In Eve, what she can’t answer becomes a durable unanswered item, and she tells the customer it is “with the support team” only once the ticket is accepted. How unanswered questions become lessons

The next operating layer

Better models raise the ceiling. Better harnesses make the work hold.

The source’s most durable idea is not that harnesses replace models. It is that intelligence becomes operational only when tools, memory, permissions, execution, observability, validation, and human oversight work as one system.

Deyaf’s builder turns that into a first step: about ten minutes to a knowledge preview and a saved setup, no account needed.

Build your assistant