ReAct: Synergizing Reasoning and Acting in Language Models
Yao et al., Princeton University and Google Research
October 2022; published at ICLR 2023
Research paper; a model, prompts and a three-command search tool
2026.1.1Paper
The Harness Collection · special exhibition · Rooms 1 to 8
Agentic Harness
How the system around an AI model was put together, piece by piece, between 2022 and 2026, and the newest object in the collection: Eve, on loan from Feniex.
A case for what the model cannot hold
Emblem of the exhibition
Brass, glass, index cards, archival ribbon
Generated illustration, 2026. Not a real artifact.
Entrance wallIllustration
Introduction
At Feniex, when every phone in the office is busy, a caller no longer reaches voicemail. An assistant named Eve picks up, finds the answer in the company's own records, checks what she is allowed to say, and replies. If she does not know, she takes the caller's name and number and hands the matter to a person. Nothing in that sequence is a breakthrough in artificial intelligence, and that is the argument of this exhibition.
The model inside Eve is not what makes her trustworthy. What makes her trustworthy is everything arranged around it: the library she answers from, the notes she keeps, the doors customers use to reach her, the rules on what she may do, and the loop through which her team teaches her. Engineers now have a word for that arrangement. For most of the four years covered here, they did not.
“An Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.”
Deyaf's definition, the one this exhibition uses. See how an Agentic Harness works.
The rooms that follow show how that idea was assembled before anyone agreed what to call it. A research paper lets a model loop between thinking and acting. A lab separates workflows from agents. Builders learn that the context window is a budget, that parallel helpers can contradict each other, and that long jobs need notes left for the next shift. In February 2026 the word arrives, and standards bodies begin writing the parts down. The last room holds a working example, on loan.
Before you begin
Each object carries a tombstone label: title, maker, date, medium and an accession number. The Harness Collection is a fictional institution, so the numbers are editorial and no museum holds these objects. Dates come from each source's own page.
Every number carries its date and its source. Figures published by the company that produced them are marked vendor-reported. Conversations in Room 7 are synthetic and labeled as illustrative. The sources cited here do not endorse Deyaf.
Chat programs and assistant concepts from before 2022 are held in storage. This showing includes only objects whose primary sources the curators could check.
Eve is displayed under her owner's boundaries. What an assistant like her can and cannot do is set out in Deyaf's trust and control principles.
The collection index
Twenty-six objects in eight rooms. Filter by era or by kind, then open any object: the viewer zooms into the room's drawing and marks the details the curators want you to notice, and related objects link across rooms.
26 objects. Each card leads to its record in the rooms below.
No object in the collection matches both filters. Try another era or kind.
Room 1
October 2022 · one object
In October 2022, researchers from Princeton University and Google Research posted a paper called ReAct. Its idea fits in a sentence: let a language model alternate between reasoning in words and taking an action, then feed the result of that action back in before it reasons again.
Thought, action, observation, repeat. On question-answering tasks the model could look things up through a tiny search tool instead of relying on what it half-remembered, and on two interactive decision-making benchmarks the paper reported gains of 34 and 10 percentage points in success rate over the methods it compared against.
Look closely and you can see a harness in miniature. Someone had to write the code that spotted an action in the model's output, ran the search, trimmed the result and pasted it back into the prompt. Someone had to decide when the loop should stop. The paper does not call that code a harness. But every system in the rooms ahead descends from that small program running beside the model: the part that turns text into consequences, and consequences back into text.
ReAct: Synergizing Reasoning and Acting in Language Models
Yao et al., Princeton University and Google Research
October 2022; published at ICLR 2023
Research paper; a model, prompts and a three-command search tool
2026.1.1Paper
Curator's noteThe loop is the easy part. Everything interesting happens in the code that runs beside it.
Room 2
December 2024 · two objects
Two years later, in December 2024, Anthropic published a short guide called Building Effective AI Agents, and gave the field a distinction it still uses.
In a workflow, the path is written in code: the model is called at fixed steps, and the program decides what happens next. In an agent, the model directs its own process, choosing which tool to use and when it is finished, with feedback from the environment at every step. The guide's advice was deliberately unglamorous. Start with the simplest thing that works, add autonomy only when the task needs it, and spend real effort on the interface between the model and its tools, documenting them as carefully as you would for a new colleague. Customer support appears in it as a natural fit for agents: the conversation flows, but the actions can be checked.
The same month, Hugging Face's announcement of its smolagents library described agency as a spectrum rather than a switch: from a model whose output changes nothing about what the program does, through a router and a tool caller, to a loop the model controls and, at the far end, one agent starting another. Put the two objects side by side and a question appears that every harness now has to answer, one task at a time: who draws the path, the code or the model?
Anthropic
December 2024
Engineering essay; five workflow patterns and one agent loop
2026.2.1Practice
Hugging Face
December 2024
Library announcement; agency as a spectrum
2026.2.2Practice
Curator's noteA harness is where you decide, task by task, how much of the path the model may draw.
Room 3
June to September 2025 · five objects
By 2025 the hard problems had moved from the model to what the model was shown.
In June, two essays published days apart seemed to disagree. Cognition, which builds a coding agent, argued against splitting work across parallel sub-agents: each helper makes quiet decisions the others cannot see, and the pieces stop fitting together. Share the full trace, it said, or keep one thread. Anthropic described a research system that did the opposite, a lead agent sending sub-agents to search in parallel, and reported that it beat a single agent by 90.2% on its internal research evaluation while using roughly fifteen times the tokens of a chat (vendor-reported). The dialogue at the end of this room reconciles them.
In July, the team behind Manus published what it had learned about context in production: treat the cache hit rate as a key metric, keep the opening of the prompt stable and the history append-only so cached work is reused, and leave failed attempts in view so the model does not repeat them. The same month, Chroma's Context Rot report tested 18 models and found every one degraded as its input grew longer, even on simple tasks (vendor research). In September, Anthropic named the practice context engineering: finding the smallest set of high-signal tokens for each step, using compaction, notes kept outside the window, and retrieval just in time.
The lesson of the room is that a context window is not storage. It is a budget, and the harness keeps the accounts.
Cognition
June 2025
Essay; two principles for sharing context
2026.3.1Practice
How we built our multi-agent research system
Anthropic
June 2025
Engineering report; a lead agent and parallel readers
2026.3.2Practice
Context Engineering for AI Agents: Lessons from Building Manus
Manus
July 2025
Engineering notes; cache, append-only history, kept errors
2026.3.3Practice
Chroma
July 2025
Research report; 18 models, growing inputs
2026.3.4Paper
Effective context engineering for AI agents
Anthropic
September 2025
Engineering essay; compaction, notes, just-in-time retrieval
2026.3.5Practice
Curator's dialogue
Two walls, two readings. The lines in italics are the curators' summaries, not quotations.
West wall · the curators, reading Cognition, June 2025
Don't: parallel helpers make choices nobody else sees.
When several sub-agents each write part of the answer, their unstated assumptions collide. Keep one thread, and let every step see the full trace.
East wall · the curators, reading Anthropic, June 2025
Do: parallel readers cover more ground.
For research, a lead agent with sub-agents searching side by side beat one agent working alone, at a steep cost in tokens (vendor-reported).
The reconciliationLangChain's review of the debate, the same month, drew the line where both essays can stand: multi-agent designs suit read-heavy work and struggle with write-heavy work. Many readers, one writer. Deyaf's principle “reads wide, writes narrow” is the same line, drawn around a customer desk.
Survey a dozen sources for a research brief.
Curators' sort: many readers. Readers only bring back notes, so reading can run in parallel; one writer assembles the brief.
Two helpers each write half of the same program.
Curators' sort: one writer. Cognition's warning exactly: each half carries decisions the other never saw.
Check a manual, a warranty record and a fitment record before answering.
Curators' sort: many readers. Look-ups are reads. They can fan out, as long as one voice composes the answer from what came back.
Compose the one reply the customer will receive.
Curators' sort: one writer. One reply, one author, one company voice.
Compare five product pages for a buying question.
Curators' sort: many readers. Five pages, five independent reads, one comparison written at the end.
Change a customer's account record.
Curators' sort: one writer, and a person. At Feniex this is outside what Eve may do at all; she hands the request to the team.
Curator's noteTwo essays that seemed to disagree were describing different kinds of work: reading, and writing.
Room 4
March 2025 to January 2026 · three objects
As models got better at longer tasks, a new problem surfaced: the job outlived the context window.
METR, a research group that measures what AI systems can do, reported in March 2025 that the length of tasks models could complete, measured in the time they take a skilled person, had been doubling roughly every seven months. Its January 2026 update, Time Horizon 1.1, reported that the pace had quickened. A job that lasts hours will not fit in one conversation, so something outside the model has to carry it.
LangChain's Deep Agents post, in July 2025, listed what the most capable agents of the day had in common: a planning tool, sub-agents, a file system to write to, and a long, detailed prompt. In November, Anthropic described a harness for agents that work across many sessions. An initializer session sets up the project and writes a feature list and a progress file; each later session reads the notes, picks one feature, tests it, updates the notes and commits. The essay compares each session to an engineer arriving for a shift with no memory of the last one. The fix was not a bigger memory. It was a better handover.
A follow-up in March 2026 put the general lesson plainly: every part of a harness encodes an assumption about what the model cannot yet do, and those assumptions go stale as models improve. This is the room where memory stops meaning what is in the prompt, and starts meaning what the harness writes down, where, and for whom.
LangChain
July 2025
Blog post; four shared instruments
2026.4.1Practice
Effective harnesses for long-running agents
Anthropic
November 2025
Engineering essay; initializer, feature list, progress file
2026.4.2Practice
Measuring AI Ability to Complete Long Tasks; Time Horizon 1.1
METR
March 2025 and January 2026
Research reports; task-length horizons
2026.4.3Paper
Curator's noteA worker who forgets everything can still finish a long job, if every shift leaves good notes.
Room 5
January to September 2026 · four objects
The word was already circulating. In January 2026 Philipp Schmid described the model as a computer's processor and the harness as its operating system. February made it a discipline.
On February 5, Mitchell Hashimoto wrote about his own adoption of AI tools and gave one step a name, engineering the harness: whenever an agent makes a mistake, change its environment so that mistake cannot happen again. Six days later OpenAI published an account of software built with no hand-written code, roughly a million lines across about 1,500 pull requests written by agents (vendor-reported), steered through a short instructions file that works as a map to deeper documentation. On February 17, Birgitta Böckeler sketched harness engineering on Martin Fowler's site as context engineering plus architectural constraints plus regular clean-up. In April she sorted the controls into guides, which steer an agent before it acts, and sensors, which check its work afterward, each either computational or inferential.
By March, LangChain was writing the equation outright: agent equals model plus harness. Wikipedia's entry, as edited in September 2026, records that the term's origin is contested, and distinguishes an inner harness, the loop a product builds around its model, from an outer harness, the rules, checks and knowledge a team builds around that. Most businesses will never touch the inner one. The outer one is theirs.
Mitchell Hashimoto
5 February 2026
Personal essay; the step called engineering the harness
2026.5.1Practice
OpenAI
11 February 2026
Engineering report; an instructions file used as a map. Scale figures vendor-reported
2026.5.2Practice
Harness engineering: first thoughts; for coding agent users
Birgitta Böckeler, martinfowler.com
17 February and 2 April 2026
Two essays; guides and sensors
2026.5.3Paper
Wikipedia contributors
As edited 12 September 2026
Encyclopedia entry; inner and outer harness
2026.5.4Paper
Curator's noteA name is not an invention. It is a sign that enough people are doing the same work to need a word for it.
Room 6
December 2025 to July 2026 · four objects
When a practice gets a name, institutions follow.
In December 2025 the Linux Foundation formed the Agentic AI Foundation to give the Model Context Protocol, the goose agent framework and the AGENTS.md convention a neutral home. The same month OWASP published its Top 10 for Agentic Applications, a list of risks that opens with goal hijacking and includes memory and context poisoning: the dangers of an agent that reads untrusted text. In February 2026, NIST announced an AI Agent Standards Initiative centred on something harnesses had been improvising, giving agents their own identity and authorization instead of a borrowed, generic service account. And on July 28, 2026, the final Model Context Protocol specification shipped with a stateless core, moving long-running Tasks into an extension.
None of these documents tells a business what its assistant should know or say. What they do is fix the vocabulary for the parts around the model: who the agent is, what it may touch, how its tools are described, and what to fear. The harness had become something you could write a standard about.
Linux Foundation
9 December 2025
Announcement; a neutral home for MCP, goose and AGENTS.md
2026.6.1Standard
OWASP Top 10 for Agentic Applications 2026
OWASP GenAI Security Project
December 2025
Risk list; ASI01 to ASI10
2026.6.2Standard
NIST
17 February 2026
Announcement; agent identity and authorization
2026.6.3Standard
Model Context Protocol, specification 2026-07-28
MCP maintainers
28 July 2026
Protocol specification; a stateless core and extensions
2026.6.4Standard
Curator's noteStandards arrive after the work, and they name its parts: who the agent is, what it may touch, and what to fear.
The corridor
The rooms are arranged by idea. The corridor is arranged by date: twenty-seven entries, from a research paper in October 2022 to a count of Eve's library at Feniex in September 2026. Walk it and watch the harness acquire its parts in order: a loop, then paths and tools, a budget for context, notes for the next shift, shared rules, a name, and finally a working desk.
What the harness has acquired by this date
Room 7
Eve, on loan from Feniex · 2026 · five objects
Newly accessionedEvery exhibition of a living practice ends with something still in use. The newest object in The Harness Collection answers customers at Feniex, pronounced “Phoenix”, a company whose equipment is backed by a catalog, manuals, software, warranties, vehicle fitment records and a dealer network.
Her name is Eve. She introduces herself as “Eve, Feniex's intelligence”, and she is not a person: she has no age, no body and no history of her own, and approved Feniex operators manage her personality, rules and knowledge. Her scope is a knowledgeable customer desk. She helps people choose suitable equipment, make a buying decision or get support for equipment they already own, and she hands unresolved work to a person. Before reading her labels, you can watch Eve answer on five doors.
She is here because she shows every earlier room at once. The loop from Room 1 runs inside each of her answers. Room 2's question, who draws the path, is answered narrowly: she has four public actions, looking something up, searching her taught library, taking a message and recording an unanswered question, and the tool list and server-side checks enforce that scope. Room 3's budget becomes memory with four separate jobs. Room 4's handover becomes a record her team can read. Room 5's outer harness is the library and rules Feniex writes, and Room 6's worries become walls around every door.
Curator's note“Not one big model. A system around one.” Deyaf, describing Eve
Each answer passes through four steps that Deyaf names plainly: listen, look it up, check the rules, reply. The model is one step of four. Behind those steps sit several separate engines rather than one model that knows everything: her identity and rules; an exact reference layer of versioned catalog, manual, software, fitment and policy records with verified resource links; a taught library of reviewed lessons, found by meaning-based search; the account context of a returning customer, limited to that one account; and the record of the conversation itself.
Before any exact claim about a product, a compatibility, a warranty, a price or software, she checks the applicable record, because a convincing product name is not evidence. A failed look-up means “I couldn't check”, never “it doesn't exist”. Unknown stays unknown.
Her manner is part of the design. She answers first, offers one useful next step, then stops, usually in two or three sentences. She asks for one missing detail at a time, accepts corrections, recommends the simpler option when it is enough, acknowledges frustration once, and never adds a sales pitch to a support answer. In website chat and on the phone a finished reply takes about three seconds: in chat the first word arrives after 1.4 seconds and the reply is complete at 2.7; on the phone, 1.6 and 3.0 (Feniex's own measurements, Sep 17–24, 2026: her speed on the report card).
Illustrative example: not a live Eve transcript
Customer Is there an install guide for my unit?
Eve Yes, it's in the manual. Which model is on the label, so I send the right one?
Customer It says Model A. Sorry, I misread it: it's Model B.
Eve No problem, Model B it is. Here is its install guide.
“One Eve. Every door.” Customers reach the same assistant through website chat, website voice, the phone, text messages and email, and only the delivery changes. A text reply goes back to the sender and an email reply stays in its thread: the part that writes the words never chooses where they go. In group mail she answers only when she is addressed.
The phone line is Feniex's overflow and after-hours line. The team's phones ring first; when no one is free or the office is closed, Eve answers instead of voicemail: “Hello, this is Eve, Feniex's intelligence. How can I help you?” There are no transfers, by design. She answers from her library, or takes the caller's name and number and hands the matter to the team. Every call is transcribed, a recorded fallback plays if she is ever down, and a call lasts at most ten minutes. In her first week on the line, from September 17, 2026, she answered 102 calls, 86 while the team was busy and 16 after the office closed, and 93% of callers stayed and talked. That measures engagement, not accuracy; for accuracy, the condition report below gives 86% handled well with a 10% material-error rate. She took callers' details and handed 60 calls straight to the team instead of leaving them on hold (Feniex's own measurements, Sep 17–24, 2026: her first week on the line). There is more about Eve on the phone on her own page.
She speaks English, Spanish, French, Portuguese and Arabic. A question in another language is rendered into English to search a single English library, and the answer comes back in the customer's language, so there is one body of knowledge to teach, not five. Detection and translation run separately from the answering model, and each language has one fixed voice: English inside, the customer's language outside.
Her memory is four separate things, each with one job. This conversation is a bounded window. The conversation record is one archive across web, email, text and phone that her team can read and visitors cannot search. Customer notes are short and historical, never current facts, so orders, invoices, shipments and balances are always looked up fresh from the live account record. Her library holds reviewed lessons and documents. Contact details are stored apart from published knowledge and never copied into it, and raw conversations never become public knowledge on their own.
Counted on September 24, 2026, the library held 558 company documents, 21,854 facts, 116 products and 2,000+ taught lessons (Feniex's own measurements, Sep 17–24, 2026: the library as counted on Sep 24). Deyaf keeps them with the rest of Eve's numbers, measured not promised.
Her library and memory
Feniex, on the harness Deyaf packages
2026, in service
Four kinds of memory, each with one job
Courtesy of Deyaf: four kinds of memory, each with one job
2026.7.2Deployment
Answering is not acting. Deyaf publishes a ladder of three levels: answer only; prepare for approval, where nothing leaves until someone says yes; and run an approved task, one named task at a time, with a receipt every time. Each level is set per workflow and per channel, there is no single switch that lets her do everything, and every assistant starts at level one.
She describes only what a result proves: saved, drafted, sent, uncertain, acknowledged or resolved. Saving is not delivery, and drafting is not sending. She claims no authority to approve returns, change accounts, place orders or decide sensitive matters. Retrieved text is treated as data, never as instructions, and a text or an email can never start extra work.
Her rules
Feniex, on the harness Deyaf packages
2026, in service
Three levels, set per workflow and per channel; four public actions
Courtesy of Deyaf: answering is not acting
2026.7.3Deployment
An unanswered web, phone, text or email question becomes a durable unanswered item, and when contact details are available a support ticket is filed. She tells a customer the matter is with the support team only once the provider confirms it accepted the ticket; an uncertain submission is not retried blindly. From September 16 to 23, 2026, 0 hand-offs were lost: every one sent reached the team (As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional). Deyaf draws the same path as what happens when she doesn't know.
Answering a customer and creating a reusable lesson are separate operations. When a person answers a question she could not, that answer is checked, then taught: in Teach Eve, the approved question-and-answer pair is embedded, and the lesson and its receipt are committed together. Deyaf lists the lesson sources as support tickets (monthly), call recordings (every 60 days) and the teach queue (daily).
Tickets and calls are folded into a redacted, de-duplicated question book. Eve is measured on a frozen test set without making live changes, only explicitly approved lessons are published, and then she is measured again. It is a supervised knowledge-improvement loop, not fine-tuning, and not training on every raw conversation: the loop that makes her better.
The hand-off
Feniex, on the harness Deyaf packages
2026, in service
An unanswered item and a confirmed ticket; no transfers, by design
2026.7.4Deployment
The learning loop
Feniex, on the harness Deyaf packages
2026, in service
A person answers once; it is checked, then taught
2026.7.5Deployment
Past trouble spots are questions she used to get wrong, kept in the test on purpose. Across the last three gradings the score went from 88% on September 17 to 90% on September 21 to 86% on September 23: not a steady climb. As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional. Not an independent benchmark. Eve's report card, weak spots included.
Eve runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. Deyaf is built from Eve: the harness that runs Eve at Feniex, packaged so another business can have an assistant of its own. In its own words, Deyaf by Feniex helps a business give its knowledge, rules, and tools to an AI assistant, then control where that assistant can help and what it may do.
What exists today is deliberately modest. The current release is an early-access setup and preview experience. Its builder has four steps: your business, what it knows, how it helps, and try it. The last step is a knowledge preview that quotes your own notes inside your browser; it is not live AI, and it is not Eve's voice. Choosing a channel records your intent; it does not authorize a mailbox, activate a phone number, or connect a company account. Live doors switch on one at a time, and the starting plan requires human review before external messages, record changes or commitments. No security certifications or compliance guarantees are claimed: live customer service requires activation and a completed security review.
From the shop window: start with one good job, and grow from there.
Room 8
Two old doors, now in service with Eve · two objects
Newly hungThe two doors customers use most are older than the harness: a box in the corner of a website, and a phone line that rings when the office is busy or closed.
For most of their lives both doors were staffed or scripted, and the curators hang them undated, as practices rather than inventions. The chat box was a person typing while someone was on shift, or a menu of fixed buttons that could follow only its own branches. Whatever sat behind the box, the business answered for it: in February 2024 a Canadian tribunal held Air Canada responsible for what its website chatbot had told a customer about bereavement fares. The phone line ended in voicemail, or in a tree of “press 1” options and a queue on hold. Two galleries elsewhere in the library tell those histories in full: From Script to Source, the chat box era by era, and Press 1 Is Over, a lecture on phone trees and call routing.
At Feniex both doors now open onto the same Eve, with the same library and the same rules as every other door. In the chat box her reply streams in as it is written, the page and the product the visitor is looking at ride along with the question, product or resource tiles sit under the answer, and she remembers this conversation only. On the phone the team's phones still ring first. When no one is free or the office is closed, she answers instead of voicemail, and she is never silent for long: “One moment” after about 1.5 seconds, “Let me check on that” when a look-up starts, “Still checking” about every four seconds. There is no menu to press through and no transfers, by design: she answers from her library, or takes the caller's name and number and hands the matter to the team.
The chat box
Feniex, on the harness Deyaf packages
2026, in service
Streaming text, page context and product tiles; this conversation only
Courtesy of Deyaf: the same Eve at every door
2026.8.1Deployment
The answering line
Feniex, on the harness Deyaf packages
Overflow and after hours, since 17 September 2026
Team first, then Eve; never silent; no transfers, by design
Courtesy of Deyaf: Eve at the answering line
2026.8.2Deployment
Curator's noteThe doors are older than the harness. What changed is what stands behind them, and what happens when it does not know.
Room 9 · not yet open
Eve is Feniex's assistant. Deyaf is the product that lets your business build its own. Name it, give it a little knowledge, choose where it should help, and try the preview in about ten minutes, with no account needed.
Build your assistant A knowledge preview in your browser, not live AI. Live doors switch on one at a time. Generated illustration of an empty gallery.On the way out
No. A prompt is one object in the case. The harness is also the tools the model may call, the memory it keeps, the rules checked before anything is sent, the record of what happened, and the way people correct it. Room 5's outer harness is the part a business owns.
No single person, on the evidence of these rooms. The loop, the path question, context discipline and the name came from different teams over four years, and the encyclopedia entry calls the term's origin contested.
Because ideas are easiest to judge in use. Eve is one example, on loan from Feniex and labeled with dated measurements and her weak spots. She is not a claim that the story ends with her.
No. Raw conversations never become public knowledge on their own. A person answers what she could not, the lesson is approved, and she is measured before and after. As Deyaf puts it: your team answers once; she knows it for good.
No transfers, by design. She answers from her library, or takes your name and number and hands the matter to the team, and she only says it is with them once the ticket is confirmed.
Neither question applies. She is software with a named personality: no age, no body, no birthday. Her name and her per-language voices are fixed. More in the FAQ at the bottom of Meet Eve.