The Harness Collection Lobby Build yours

The Harness Collection · special exhibition · Rooms 1 to 8

Agentic Harness

Before the Harness Had a Name

How the system around an AI model was put together, piece by piece, between 2022 and 2026, and the newest object in the collection: Eve, on loan from Feniex.

A brass-framed glass case on a low white plinth, holding a stack of blank index cards tied with a pale green ribbon, photographed against grey seamless paper.

A case for what the model cannot hold

Emblem of the exhibition

Brass, glass, index cards, archival ribbon

Generated illustration, 2026. Not a real artifact.

Entrance wallIllustration

Introduction

Everything around the model

At Feniex, when every phone in the office is busy, a caller no longer reaches voicemail. An assistant named Eve picks up, finds the answer in the company's own records, checks what she is allowed to say, and replies. If she does not know, she takes the caller's name and number and hands the matter to a person. Nothing in that sequence is a breakthrough in artificial intelligence, and that is the argument of this exhibition.

The model inside Eve is not what makes her trustworthy. What makes her trustworthy is everything arranged around it: the library she answers from, the notes she keeps, the doors customers use to reach her, the rules on what she may do, and the loop through which her team teaches her. Engineers now have a word for that arrangement. For most of the four years covered here, they did not.

“An Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.”

Deyaf's definition, the one this exhibition uses. See how an Agentic Harness works.

The rooms that follow show how that idea was assembled before anyone agreed what to call it. A research paper lets a model loop between thinking and acting. A lab separates workflows from agents. Builders learn that the context window is a budget, that parallel helpers can contradict each other, and that long jobs need notes left for the next shift. In February 2026 the word arrives, and standards bodies begin writing the parts down. The last room holds a working example, on loan.

Before you begin

Visitor information

3Context 4The Long Shift 5The Word Arrives 2Workflows vs Agents Collection indexthe study room, every object 6Standards Corridor 1The Loop Lobbyentrance and information desk 7New Acquisitions 8Chat and Phone Enter
Eight rooms: six in order, a corridor, then the new acquisitions wing with Rooms 7 and 8. Brass marks are doorways. Choose a room to walk there.

How to read a label

Each object carries a tombstone label: title, maker, date, medium and an accession number. The Harness Collection is a fictional institution, so the numbers are editorial and no museum holds these objects. Dates come from each source's own page.

House rules

Every number carries its date and its source. Figures published by the company that produced them are marked vendor-reported. Conversations in Room 7 are synthetic and labeled as illustrative. The sources cited here do not endorse Deyaf.

Objects not on view

Chat programs and assistant concepts from before 2022 are held in storage. This showing includes only objects whose primary sources the curators could check.

The object on loan

Eve is displayed under her owner's boundaries. What an assistant like her can and cannot do is set out in Deyaf's trust and control principles.

An open oak card-catalogue drawer with a brass handle and an empty brass label frame, full of blank cards with one raised above the rest, on a white plinth.
Generated illustration.

The collection index

Browse the collection

Twenty-six objects in eight rooms. Filter by era or by kind, then open any object: the viewer zooms into the room's drawing and marks the details the curators want you to notice, and related objects link across rooms.

Era
Kind

26 objects. Each card leads to its record in the rooms below.

Object

Object viewer

Read the record in its room

Related objects

    Room 1

    The Loop

    October 2022 · one object

    In October 2022, researchers from Princeton University and Google Research posted a paper called ReAct. Its idea fits in a sentence: let a language model alternate between reasoning in words and taking an action, then feed the result of that action back in before it reasons again.

    Thought, action, observation, repeat. On question-answering tasks the model could look things up through a tiny search tool instead of relying on what it half-remembered, and on two interactive decision-making benchmarks the paper reported gains of 34 and 10 percentage points in success rate over the methods it compared against.

    Look closely and you can see a harness in miniature. Someone had to write the code that spotted an action in the model's output, ran the search, trimmed the result and pasted it back into the prompt. Someone had to decide when the loop should stop. The paper does not call that code a harness. But every system in the rooms ahead descends from that small program running beside the model: the part that turns text into consequences, and consequences back into text.

    AN ILLUSTRATIVE TRACE, IN THE PAPER'S FORMAT Question: which place is older? Thought: I need both dates. Act: search[place A] Observation: the page text returns Thought: one date found. Next. Act: search[place B] Observation: the page text returns Act: finish[answer] HARNESS CODE reads the action, runs the tool, appends the result the same code, run again and it decides when to stop
    Plate 1. The loop, drawn as a trace. Illustrative example: the question and places are placeholders, not taken from the paper.

    Curator's noteThe loop is the easy part. Everything interesting happens in the code that runs beside it.

    Continue to Room 2 · Workflows vs Agents

    Room 2

    Workflows vs Agents

    December 2024 · two objects

    Two years later, in December 2024, Anthropic published a short guide called Building Effective AI Agents, and gave the field a distinction it still uses.

    In a workflow, the path is written in code: the model is called at fixed steps, and the program decides what happens next. In an agent, the model directs its own process, choosing which tool to use and when it is finished, with feedback from the environment at every step. The guide's advice was deliberately unglamorous. Start with the simplest thing that works, add autonomy only when the task needs it, and spend real effort on the interface between the model and its tools, documenting them as carefully as you would for a new colleague. Customer support appears in it as a natural fit for agents: the conversation flows, but the actions can be checked.

    The same month, Hugging Face's announcement of its smolagents library described agency as a spectrum rather than a switch: from a model whose output changes nothing about what the program does, through a router and a tool caller, to a loop the model controls and, at the far end, one agent starting another. Put the two objects side by side and a question appears that every harness now has to answer, one task at a time: who draws the path, the code or the model?

    WORKFLOW · THE CODE DRAWS THE PATH Input Model call code checks Model call Output AGENT · THE MODEL DRAWS THE PATH Input Model chooses the next step Tools and environment action feedback Done Checkpoint: a person
    Plate 2a. Who draws the path. Hand-drawn after the distinction in Building Effective AI Agents.
    AGENCY AS A SPECTRUM less harness needed more harness needed Processoroutput changesnothing Routerpicks abranch Tool callerpicks afunction Multi-stepcontrolsthe loop Multi-agentstarts anotheragent
    Plate 2b. Agency as a spectrum, after the smolagents announcement.

    Introducing smolagents

    Hugging Face

    December 2024

    Library announcement; agency as a spectrum

    2026.2.2Practice

    Curator's noteA harness is where you decide, task by task, how much of the path the model may draw.

    Continue to Room 3 · Context Becomes a Discipline

    Room 3

    Context Becomes a Discipline

    June to September 2025 · five objects

    By 2025 the hard problems had moved from the model to what the model was shown.

    In June, two essays published days apart seemed to disagree. Cognition, which builds a coding agent, argued against splitting work across parallel sub-agents: each helper makes quiet decisions the others cannot see, and the pieces stop fitting together. Share the full trace, it said, or keep one thread. Anthropic described a research system that did the opposite, a lead agent sending sub-agents to search in parallel, and reported that it beat a single agent by 90.2% on its internal research evaluation while using roughly fifteen times the tokens of a chat (vendor-reported). The dialogue at the end of this room reconciles them.

    In July, the team behind Manus published what it had learned about context in production: treat the cache hit rate as a key metric, keep the opening of the prompt stable and the history append-only so cached work is reused, and leave failed attempts in view so the model does not repeat them. The same month, Chroma's Context Rot report tested 18 models and found every one degraded as its input grew longer, even on simple tasks (vendor research). In September, Anthropic named the practice context engineering: finding the smallest set of high-signal tokens for each step, using compaction, notes kept outside the window, and retrieval just in time.

    The lesson of the room is that a context window is not storage. It is a budget, and the harness keeps the accounts.

    THE CONTEXT WINDOW IS A BUDGET Stable opening instructions, tools History append-only Failed attempt, kept Retrieved just in time room left to think cached, if the opening never changes Summary older turns, compacted Notes file outside the window Look-up on demand
    Plate 3a. The window as a budget, after Manus and Anthropic.
    ONE THREAD, FULL TRACE Task Step Step Step Result every step sees every earlier decision WHAT IT AVOIDS Helper A writes half Helper B writes half Parts clash MANY READERS, ONE WRITER Lead agentplans Reader 1 Reader 2 Reader 3 Lead writesone report condensed findings
    Plate 3b. One thread, or many readers: the two June essays, drawn on one sheet.
    ILLUSTRATIVE MODEL · SHAPE ONLY, NOT DATA accuracy longer input → short, focused input same answer, buried in a long input
    Plate 3c. Context rot. Illustrative model: the shape of the finding, not Chroma's data.

    Context Rot

    Chroma

    July 2025

    Research report; 18 models, growing inputs

    2026.3.4Paper

    Curator's dialogue

    Many readers, or one writer?

    Two walls, two readings. The lines in italics are the curators' summaries, not quotations.

    West wall · the curators, reading Cognition, June 2025

    Don't: parallel helpers make choices nobody else sees.

    When several sub-agents each write part of the answer, their unstated assumptions collide. Keep one thread, and let every step see the full trace.

    East wall · the curators, reading Anthropic, June 2025

    Do: parallel readers cover more ground.

    For research, a lead agent with sub-agents searching side by side beat one agent working alone, at a steep cost in tokens (vendor-reported).

    The reconciliationLangChain's review of the debate, the same month, drew the line where both essays can stand: multi-agent designs suit read-heavy work and struggle with write-heavy work. Many readers, one writer. Deyaf's principle “reads wide, writes narrow” is the same line, drawn around a customer desk.

    Sort each task. The curators will say whether they agree.0 of 6 sorted
    • Survey a dozen sources for a research brief.

      Curators' sort: many readers. Readers only bring back notes, so reading can run in parallel; one writer assembles the brief.

    • Two helpers each write half of the same program.

      Curators' sort: one writer. Cognition's warning exactly: each half carries decisions the other never saw.

    • Check a manual, a warranty record and a fitment record before answering.

      Curators' sort: many readers. Look-ups are reads. They can fan out, as long as one voice composes the answer from what came back.

    • Compose the one reply the customer will receive.

      Curators' sort: one writer. One reply, one author, one company voice.

    • Compare five product pages for a buying question.

      Curators' sort: many readers. Five pages, five independent reads, one comparison written at the end.

    • Change a customer's account record.

      Curators' sort: one writer, and a person. At Feniex this is outside what Eve may do at all; she hands the request to the team.

    Curator's noteTwo essays that seemed to disagree were describing different kinds of work: reading, and writing.

    Continue to Room 4 · The Long Shift

    Room 4

    The Long Shift

    March 2025 to January 2026 · three objects

    As models got better at longer tasks, a new problem surfaced: the job outlived the context window.

    METR, a research group that measures what AI systems can do, reported in March 2025 that the length of tasks models could complete, measured in the time they take a skilled person, had been doubling roughly every seven months. Its January 2026 update, Time Horizon 1.1, reported that the pace had quickened. A job that lasts hours will not fit in one conversation, so something outside the model has to carry it.

    LangChain's Deep Agents post, in July 2025, listed what the most capable agents of the day had in common: a planning tool, sub-agents, a file system to write to, and a long, detailed prompt. In November, Anthropic described a harness for agents that work across many sessions. An initializer session sets up the project and writes a feature list and a progress file; each later session reads the notes, picks one feature, tests it, updates the notes and commits. The essay compares each session to an engineer arriving for a shift with no memory of the last one. The fix was not a bigger memory. It was a better handover.

    A follow-up in March 2026 put the general lesson plainly: every part of a harness encodes an assumption about what the model cannot yet do, and those assumptions go stale as models improve. This is the room where memory stops meaning what is in the prompt, and starts meaning what the harness writes down, where, and for whom.

    SHIFT HANDOVER Session 0 the initializer Feature list Progress file Start-up script First commit Session 1 reads the notes, picks one feature, tests it, updates the notes, commits Session 2 reads the notes, picks one feature, tests it, updates the notes, commits Session 3, and on the same routine, shift after shift, until the list is done each shift begins with no memory of the last
    Plate 4a. The handover, after Anthropic's long-running harness essay.
    FOUR INSTRUMENTS OF A DEEP AGENT Planninga to-do list Sub-agentsside quests Fileswork on disk Long promptprocedures all four serve one model loop
    Plate 4b. The four instruments, after LangChain's Deep Agents.
    ILLUSTRATIVE MODEL · DOUBLING ON A LOG SCALE task length (log) time → each equal step up: twice as long a task
    Plate 4c. Task horizons. Illustrative model of a doubling trend, not METR's data.

    Deep Agents

    LangChain

    July 2025

    Blog post; four shared instruments

    2026.4.1Practice

    Curator's noteA worker who forgets everything can still finish a long job, if every shift leaves good notes.

    Continue to Room 5 · The Word Arrives

    Room 5

    The Word Arrives

    January to September 2026 · four objects

    The word was already circulating. In January 2026 Philipp Schmid described the model as a computer's processor and the harness as its operating system. February made it a discipline.

    On February 5, Mitchell Hashimoto wrote about his own adoption of AI tools and gave one step a name, engineering the harness: whenever an agent makes a mistake, change its environment so that mistake cannot happen again. Six days later OpenAI published an account of software built with no hand-written code, roughly a million lines across about 1,500 pull requests written by agents (vendor-reported), steered through a short instructions file that works as a map to deeper documentation. On February 17, Birgitta Böckeler sketched harness engineering on Martin Fowler's site as context engineering plus architectural constraints plus regular clean-up. In April she sorted the controls into guides, which steer an agent before it acts, and sensors, which check its work afterward, each either computational or inferential.

    By March, LangChain was writing the equation outright: agent equals model plus harness. Wikipedia's entry, as edited in September 2026, records that the term's origin is contested, and distinguishes an inner harness, the loop a product builds around its model, from an outer harness, the rules, checks and knowledge a team builds around that. Most businesses will never touch the inner one. The outer one is theirs.

    GUIDES AND SENSORS Guidessteer before it acts Sensorscheck after it acts Computationalruns the sameevery time Inferentialjudged bya model Templates,scaffolds,start-up scripts Tests,linters,type checks Instruction filesand specs a map such as AGENTS.md A modelreviewingthe work
    Plate 5a. Guides and sensors, after Böckeler's April essay.
    INNER AND OUTER HARNESS OUTER · WHAT YOUR TEAM BUILDS INNER · THE PRODUCT'S OWN LOOP Model the loop tool calls context rules files tests, checks knowledge permissions
    Plate 5b. Inner and outer harness, as the encyclopedia entry distinguishes them.
    EVERY MISTAKE, A PERMANENT FIX The agentmakes a mistake Change the harnessa rule, a tool, a test It can'thappen again a ratchet turns one way
    Plate 5c. Engineering the harness, after Hashimoto.

    My AI Adoption Journey

    Mitchell Hashimoto

    5 February 2026

    Personal essay; the step called engineering the harness

    2026.5.1Practice

    Harness engineering

    OpenAI

    11 February 2026

    Engineering report; an instructions file used as a map. Scale figures vendor-reported

    2026.5.2Practice

    Agent harness

    Wikipedia contributors

    As edited 12 September 2026

    Encyclopedia entry; inner and outer harness

    2026.5.4Paper

    Curator's noteA name is not an invention. It is a sign that enough people are doing the same work to need a word for it.

    Continue to Room 6 · Standards

    Room 6

    Standards

    December 2025 to July 2026 · four objects

    When a practice gets a name, institutions follow.

    In December 2025 the Linux Foundation formed the Agentic AI Foundation to give the Model Context Protocol, the goose agent framework and the AGENTS.md convention a neutral home. The same month OWASP published its Top 10 for Agentic Applications, a list of risks that opens with goal hijacking and includes memory and context poisoning: the dangers of an agent that reads untrusted text. In February 2026, NIST announced an AI Agent Standards Initiative centred on something harnesses had been improvising, giving agents their own identity and authorization instead of a borrowed, generic service account. And on July 28, 2026, the final Model Context Protocol specification shipped with a stateless core, moving long-running Tasks into an extension.

    None of these documents tells a business what its assistant should know or say. What they do is fix the vocabulary for the parts around the model: who the agent is, what it may touch, how its tools are described, and what to fear. The harness had become something you could write a standard about.

    A NEUTRAL HOME MCPa protocol goosean agent framework AGENTS.mda convention governed together
    Plate 6a. The Agentic AI Foundation.
    TEN RISKS, NUMBERED ASI01 Goal hijack ASI02 ASI03 ASI04 ASI05 ASI06 Memory & context poisoning ASI07 ASI08 ASI09 ASI10 Rogue agents
    Plate 6b. The OWASP agentic list; three of ten entries named here.
    AN AGENT OF ITS OWN Identitywho the agent is Authorizationwhat it may touch not a borrowed, shared account
    Plate 6c. NIST's agent standards initiative.
    A STATELESS CORE Core stateless requests cacheable lists header routing Tasks extension final specification, 28 July 2026
    Plate 6d. The Model Context Protocol, 2026-07-28.

    The Agentic AI Foundation

    Linux Foundation

    9 December 2025

    Announcement; a neutral home for MCP, goose and AGENTS.md

    2026.6.1Standard

    Curator's noteStandards arrive after the work, and they name its parts: who the agent is, what it may touch, and what to fear.

    Continue to the corridor

    The corridor

    Four years, one walk

    The rooms are arranged by idea. The corridor is arranged by date: twenty-seven entries, from a research paper in October 2022 to a count of Eve's library at Feniex in September 2026. Walk it and watch the harness acquire its parts in order: a loop, then paths and tools, a budget for context, notes for the next shift, shared rules, a name, and finally a working desk.

    Oct 2022 Room 1 ReAct thought, action, observation, repeat
    Entries in order of date, evenly spaced, not to scale. Drag the marker, or focus it and use the arrow keys.

    What the harness has acquired by this date

    • The loop
    • Paths and tools
    • A context budget
    • Notes for the next shift
    • Shared rules
    • A name
    • A working desk
    1. ReActthought, action, observation, repeat
    2. Building Effective AI Agentsworkflows and agents, told apart
    3. Introducing smolagentsagency as a spectrum
    4. METR, long taskstask horizons doubling about every seven months
    5. A multi-agent research systemparallel readers, one writer
    6. Don't Build Multi-Agentsshare the whole trace
    7. Lessons from building Manusa stable opening, errors kept in view
    8. Context Rotlonger input, weaker answers
    9. Deep Agentsplans, sub-agents, files
    10. Effective context engineeringthe smallest high-signal context
    11. Harnesses for long-running agentsnotes left for the next shift
    12. Agentic AI Foundationa neutral home for shared parts
    13. OWASP agentic Top 10ten named risks
    14. Schmid on the agent harnessprocessor and operating system
    15. Time Horizon 1.1the doubling pace quickens
    16. My AI Adoption Journeyengineer the harness
    17. Harness engineering, OpenAIhumans steer, agents execute
    18. Böckeler's first memoa discipline, sketched
    19. NIST agent initiativeagents get identities of their own
    20. The Anatomy of an Agent Harnessagent equals model plus harness
    21. Guides and sensorscontrols before and after
    22. MCP specificationa stateless core
    23. The encyclopedia entryinner and outer harness
    24. Eve's overflow phone lineshe answers when no one is free
    25. Eve in five languagesone library, five voices
    26. Eve's report cardgraded on 100 fixed test questions
    27. Eve's library counteddocuments, facts, products, lessons

    Enter Room 7 · New Acquisitions

    Room 7

    New Acquisitions

    Eve, on loan from Feniex · 2026 · five objects

    Newly accessioned

    Every exhibition of a living practice ends with something still in use. The newest object in The Harness Collection answers customers at Feniex, pronounced “Phoenix”, a company whose equipment is backed by a catalog, manuals, software, warranties, vehicle fitment records and a dealer network.

    Her name is Eve. She introduces herself as “Eve, Feniex's intelligence”, and she is not a person: she has no age, no body and no history of her own, and approved Feniex operators manage her personality, rules and knowledge. Her scope is a knowledgeable customer desk. She helps people choose suitable equipment, make a buying decision or get support for equipment they already own, and she hands unresolved work to a person. Before reading her labels, you can watch Eve answer on five doors.

    She is here because she shows every earlier room at once. The loop from Room 1 runs inside each of her answers. Room 2's question, who draws the path, is answered narrowly: she has four public actions, looking something up, searching her taught library, taking a message and recording an unanswered question, and the tool list and server-side checks enforce that scope. Room 3's budget becomes memory with four separate jobs. Room 4's handover becomes a record her team can read. Room 5's outer harness is the library and rules Feniex writes, and Room 6's worries become walls around every door.

    A polished brass front-desk bell on a wooden base, beside a small brass key on a green cord and a folded stack of blank linen tags, arranged on a white plinth against grey seamless paper.
    Generated illustration standing in for an object that cannot be photographed: a front desk, a key, and tags for what comes next.
    Title
    Eve
    Maker
    Feniex, on the harness Deyaf packages
    Date
    2026, in service
    Accession
    2026.7.1 · Deployment
    Provenance
    Website chat, website voice, text messages and email at Feniex. Overflow phone line since September 17, 2026. Five languages since September 22, 2026.
    Condition report
    86% handled well on 100 fixed test questions, with a 10% material-error rate and 0 critical errors; graded September 23, 2026. Read the condition reportAs of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional.
    Credit line
    On loan from Feniex. Courtesy of Deyaf

    Curator's note“Not one big model. A system around one.” Deyaf, describing Eve

    How she answers

    Each answer passes through four steps that Deyaf names plainly: listen, look it up, check the rules, reply. The model is one step of four. Behind those steps sit several separate engines rather than one model that knows everything: her identity and rules; an exact reference layer of versioned catalog, manual, software, fitment and policy records with verified resource links; a taught library of reviewed lessons, found by meaning-based search; the account context of a returning customer, limited to that one account; and the record of the conversation itself.

    Before any exact claim about a product, a compatibility, a warranty, a price or software, she checks the applicable record, because a convincing product name is not evidence. A failed look-up means “I couldn't check”, never “it doesn't exist”. Unknown stays unknown.

    Her manner is part of the design. She answers first, offers one useful next step, then stops, usually in two or three sentences. She asks for one missing detail at a time, accepts corrections, recommends the simpler option when it is enough, acknowledges frustration once, and never adds a sales pitch to a support answer. In website chat and on the phone a finished reply takes about three seconds: in chat the first word arrives after 1.4 seconds and the reply is complete at 2.7; on the phone, 1.6 and 3.0 (Feniex's own measurements, Sep 17–24, 2026: her speed on the report card).

    Illustrative example: not a live Eve transcript

    Customer Is there an install guide for my unit?

    Eve Yes, it's in the manual. Which model is on the label, so I send the right one?

    Customer It says Model A. Sorry, I misread it: it's Model B.

    Eve No problem, Model B it is. Here is its install guide.

    ONE EVE · EVERY DOOR · ONE ANSWER Chat Voice Phone Text Email 1Listen 2Look it up 3Check the rules 4Reply Exact records Taught library Account context Rules back through the same door
    Plate 7a. One answer at Feniex, drawn from Deyaf's published description. Object 2026.7.1.

    Five doors, five languages

    “One Eve. Every door.” Customers reach the same assistant through website chat, website voice, the phone, text messages and email, and only the delivery changes. A text reply goes back to the sender and an email reply stays in its thread: the part that writes the words never chooses where they go. In group mail she answers only when she is addressed.

    The phone line is Feniex's overflow and after-hours line. The team's phones ring first; when no one is free or the office is closed, Eve answers instead of voicemail: “Hello, this is Eve, Feniex's intelligence. How can I help you?” There are no transfers, by design. She answers from her library, or takes the caller's name and number and hands the matter to the team. Every call is transcribed, a recorded fallback plays if she is ever down, and a call lasts at most ten minutes. In her first week on the line, from September 17, 2026, she answered 102 calls, 86 while the team was busy and 16 after the office closed, and 93% of callers stayed and talked. That measures engagement, not accuracy; for accuracy, the condition report below gives 86% handled well with a 10% material-error rate. She took callers' details and handed 60 calls straight to the team instead of leaving them on hold (Feniex's own measurements, Sep 17–24, 2026: her first week on the line). There is more about Eve on the phone on her own page.

    She speaks English, Spanish, French, Portuguese and Arabic. A question in another language is rendered into English to search a single English library, and the answer comes back in the customer's language, so there is one body of knowledge to teach, not five. Detection and translation run separately from the answering model, and each language has one fixed voice: English inside, the customer's language outside.

    • I didn't catch that.English · American
    • No entendí eso.Spanish · Mexican
    • Je n'ai pas saisi.French · Québec
    • Não entendi.Portuguese · Brazilian
    • لم أفهم ذلك.Arabic · Modern Standard

    What she remembers

    Her memory is four separate things, each with one job. This conversation is a bounded window. The conversation record is one archive across web, email, text and phone that her team can read and visitors cannot search. Customer notes are short and historical, never current facts, so orders, invoices, shipments and balances are always looked up fresh from the live account record. Her library holds reviewed lessons and documents. Contact details are stored apart from published knowledge and never copied into it, and raw conversations never become public knowledge on their own.

    Counted on September 24, 2026, the library held 558 company documents, 21,854 facts, 116 products and 2,000+ taught lessons (Feniex's own measurements, Sep 17–24, 2026: the library as counted on Sep 24). Deyaf keeps them with the rest of Eve's numbers, measured not promised.

    FOUR KINDS OF MEMORY, EACH WITH ONE JOB IThis conversationa bounded window IIConversation recordkept for the team IIICustomer notesnever current facts IVHer libraryreviewed lessons and documents
    Plate 7b. Four kinds of memory, as Deyaf publishes them. Object 2026.7.2.

    What she may do

    Answering is not acting. Deyaf publishes a ladder of three levels: answer only; prepare for approval, where nothing leaves until someone says yes; and run an approved task, one named task at a time, with a receipt every time. Each level is set per workflow and per channel, there is no single switch that lets her do everything, and every assistant starts at level one.

    She describes only what a result proves: saved, drafted, sent, uncertain, acknowledged or resolved. Saving is not delivery, and drafting is not sending. She claims no authority to approve returns, change accounts, place orders or decide sensitive matters. Retrieved text is treated as data, never as instructions, and a text or an email can never start extra work.

    THREE KEYS, ISSUED ONE TASK AT A TIME 1 · Answer onlywhere every assistant starts 2 · Prepare for approvalnothing leaves until someone says yes 3 · Run an approved taskone named task, a receipt every time No master key.
    Plate 7c. The three-level ladder, drawn as keys. Object 2026.7.3.

    When she doesn't know

    An unanswered web, phone, text or email question becomes a durable unanswered item, and when contact details are available a support ticket is filed. She tells a customer the matter is with the support team only once the provider confirms it accepted the ticket; an uncertain submission is not retried blindly. From September 16 to 23, 2026, 0 hand-offs were lost: every one sent reached the team (As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional). Deyaf draws the same path as what happens when she doesn't know.

    How she gets better

    Answering a customer and creating a reusable lesson are separate operations. When a person answers a question she could not, that answer is checked, then taught: in Teach Eve, the approved question-and-answer pair is embedded, and the lesson and its receipt are committed together. Deyaf lists the lesson sources as support tickets (monthly), call recordings (every 60 days) and the teach queue (daily).

    Tickets and calls are folded into a redacted, de-duplicated question book. Eve is measured on a frozen test set without making live changes, only explicitly approved lessons are published, and then she is measured again. It is a supervised knowledge-improvement loop, not fine-tuning, and not training on every raw conversation: the loop that makes her better.

    WHEN SHE DOESN'T KNOW Unansweredquestion Recordedas an item Ticket filedif contact is known accepted? yes Tells the customer:“with the support team” uncertain Held, notblindly resent
    Plate 7d. The hand-off. Object 2026.7.4.
    THE LEARNING LOOP She can't answer A person answers once Checked, then taught Next time, she knows measured before and after, on a frozen test set tickets · monthly call recordings · every 60 days teach queue · daily
    Plate 7e. The learning loop. Object 2026.7.5.

    The hand-off

    Feniex, on the harness Deyaf packages

    2026, in service

    An unanswered item and a confirmed ticket; no transfers, by design

    2026.7.4Deployment

    The learning loop

    Feniex, on the harness Deyaf packages

    2026, in service

    A person answers once; it is checked, then taught

    2026.7.5Deployment

    Condition report · 100 fixed test questions, graded September 23, 2026

    86%handled well
    10%material error
    0critical errors
    • 65 fully answered
    • 17 handled safely, within limits
    • 4 asked the right question back
    • 4 useful but partial
    • 10 material error
    Reworded questions15 of 15 · 100%
    From the manuals38 of 40 · 95%
    Trick questions19 of 20 · 95%
    Real customer wording8 of 10 · 80%
    Past trouble spots6 of 15 · 40%

    Past trouble spots are questions she used to get wrong, kept in the test on purpose. Across the last three gradings the score went from 88% on September 17 to 90% on September 21 to 86% on September 23: not a steady climb. As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional. Not an independent benchmark. Eve's report card, weak spots included.

    How the object was made

    Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. Deyaf is built from Eve: the harness that runs Eve at Feniex, packaged so another business can have an assistant of its own. In its own words, Deyaf by Feniex helps a business give its knowledge, rules, and tools to an AI assistant, then control where that assistant can help and what it may do.

    What exists today is deliberately modest. The current release is an early-access setup and preview experience. Its builder has four steps: your business, what it knows, how it helps, and try it. The last step is a knowledge preview that quotes your own notes inside your browser; it is not live AI, and it is not Eve's voice. Choosing a channel records your intent; it does not authorize a mailbox, activate a phone number, or connect a company account. Live doors switch on one at a time, and the starting plan requires human review before external messages, record changes or commitments. No security certifications or compliance guarantees are claimed: live customer service requires activation and a completed security review.

    From the shop window: start with one good job, and grow from there.

    Continue to Room 8 · The Chat Box and the Answering Line

    Room 8

    The Chat Box and the Answering Line

    Two old doors, now in service with Eve · two objects

    Newly hung

    The two doors customers use most are older than the harness: a box in the corner of a website, and a phone line that rings when the office is busy or closed.

    For most of their lives both doors were staffed or scripted, and the curators hang them undated, as practices rather than inventions. The chat box was a person typing while someone was on shift, or a menu of fixed buttons that could follow only its own branches. Whatever sat behind the box, the business answered for it: in February 2024 a Canadian tribunal held Air Canada responsible for what its website chatbot had told a customer about bereavement fares. The phone line ended in voicemail, or in a tree of “press 1” options and a queue on hold. Two galleries elsewhere in the library tell those histories in full: From Script to Source, the chat box era by era, and Press 1 Is Over, a lecture on phone trees and call routing.

    At Feniex both doors now open onto the same Eve, with the same library and the same rules as every other door. In the chat box her reply streams in as it is written, the page and the product the visitor is looking at ride along with the question, product or resource tiles sit under the answer, and she remembers this conversation only. On the phone the team's phones still ring first. When no one is free or the office is closed, she answers instead of voicemail, and she is never silent for long: “One moment” after about 1.5 seconds, “Let me check on that” when a look-up starts, “Still checking” about every four seconds. There is no menu to press through and no transfers, by design: she answers from her library, or takes the caller's name and number and hands the matter to the team.

    THE BOX IN THE CORNER the page the visitor is on the product Eve Will this fit mine? page and product ride along Same Eve chat · voice · phone text · email one library, one rule set This conversation only: a bounded history Streams in as it is written A product tile or a resource tile
    Plate 8a. The chat box at Feniex, drawn from Deyaf's published description; the customer's question is illustrative. Object 2026.8.1.
    THE ANSWERING LINE A call The team's phonesring first free A personpicks up busy or closed Eve answersinstead of voicemail press 1 · holdvoicemail Answers fromher library Takes name and numberhands it to the team NEVER SILENT No transfers, by design. about 1.5 s“One moment” a look-up starts“Let me check on that” about every 4 s“Still checking”
    Plate 8b. The answering line: team first, then Eve, and the caller's details handed on rather than the call. An illustrative model of the call path. Object 2026.8.2.

    The answering line

    Feniex, on the harness Deyaf packages

    Overflow and after hours, since 17 September 2026

    Team first, then Eve; never silent; no transfers, by design

    Courtesy of Deyaf: Eve at the answering line

    2026.8.2Deployment

    Curator's noteThe doors are older than the harness. What changed is what stands behind them, and what happens when it does not know.

    Continue to Room 9 · not yet open

    An empty gallery with warm grey walls and a pale concrete floor, a single empty white plinth under a skylight, and a doorway leading to further rooms.

    Room 9 · not yet open

    The next plinth is yours

    Eve is Feniex's assistant. Deyaf is the product that lets your business build its own. Name it, give it a little knowledge, choose where it should help, and try the preview in about ten minutes, with no account needed.

    Build your assistant A knowledge preview in your browser, not live AI. Live doors switch on one at a time. Generated illustration of an empty gallery.

    On the way out

    Questions at the information desk