A short history of the website chat box

Agentic HarnessFrom Script to Source

The little box in the corner of a website has been rebuilt three times: first a person, then a script, then a model that could talk about anything, true or not. This is the story of the fourth box, the one that answers from sources.

A 1-bit dithered drawing of a compact desktop computer on a wooden desk, a small window open on its screen, with a coffee mug on one side and a floppy disk on the other.
Mood illustration: a small computer on a desk.

Read me first

Open almost any business website and there is a small box waiting in the corner. It has been there, in one form or another, for most of the commercial web’s life, and it has been rebuilt from the inside at least three times. First a person sat behind it. Then a script. Then a language model that could talk about anything, including things that were not true.

This page is a short history of that box, told the way the box itself might tell it: as a desktop you can poke around. Each era gets a window, and each window runs the same customer question, so you can watch what changed and what broke. The history is deliberately hedged. Dates appear only where a public source pins them down, and every conversation you see is an illustrative example written for this page, not a transcript.

The fourth window is the one this library is about. An Agentic Harness chat box looks like the other three from the outside: the same box, the same cursor, the same text arriving word by word. The difference is behind the glass. It answers from sources it can check, it knows which page you are on, it keeps the conversation bounded, and when it doesn’t know, it says so and hands the question to a person. Getting there meant keeping what each earlier era did right and fixing what each one did wrong.

The working example is EVE, Feniex’s assistant, who answers in the chat box on Feniex’s website and at every other door the company has. She arrives near the end. First, the question.

A 1-bit dithered drawing of five overlapping desktop windows with striped title bars, each holding a different empty speech bubble: rounded, square, dome-shaped, cloud and starburst.
Every era put a different voice in the same window.

One question, four boxes

Each window below is the chat box of a different era, answering the same customer. Tab to a window and press Enter (or its Run button) to step through the conversation. Drag a title bar to rearrange the desk, or focus a window’s ⇕ handle and use the arrow keys. All four conversations are shown in full. Illustrative examples for a fictional store; the policy numbers are invented for the demo.

The customer asks: “I bought this 40 days ago and it stopped working. Can I get it replaced?” (The store’s real rules: returns within 30 days; defects covered for one year.)

Era 1 · People.chat

Live chat, staffed by people · 9:40 pm, Sunday

  1. Visitor

    I bought this 40 days ago and it stopped working. Can I get it replaced?

  2. All of our agents are busy. You are number 6 in line.

  3. Our team is offline now. Leave a message and we’ll reply within one business day.

✗ Right people, wrong hour. Nobody was there to answer.

Ready
Era 2 · Script.bot

Scripted decision-tree bot

  1. Bot

    Hi! What can I help with?

    OrdersReturnsShipping
  2. Visitor

    I bought this 40 days ago and it stopped working. Can I get it replaced?

  3. Bot

    Sorry, I didn’t understand that. Please choose a topic.

  4. Visitor taps

    Returns

  5. Bot

    Returns are accepted within 30 days of purchase. Was this helpful?

    YesNo

✗ Never wrong, never helpful. The warranty path wasn’t on the tree.

Ready
Era 3 · Fluent.model

A language model with nothing around it

  1. Visitor

    I bought this 40 days ago and it stopped working. Can I get it replaced?

  2. Model

    So sorry about that! Everything we sell is covered for 90 days, no questions asked. I’ve gone ahead and approved your replacement, and it ships tomorrow.

✗ Fluent and invented. There is no 90-day rule, and nothing was approved or shipped.

Ready
Era 4 · Harness.chat

An Agentic Harness chat box

  1. Visitor

    I bought this 40 days ago and it stopped working. Can I get it replaced?

  2. Checked: warranty policy, current version · the product on this page
  3. Chat box · AI assistant

    Returns close at 30 days, but this item’s warranty covers defects for one year, so a replacement request fits. I can’t approve it myself. What’s your order number? I’ll check the order and pass it to the support team.

✓ Answered from the record, offered one next step, then stopped. No promise it couldn’t keep.

Ready

What to notice: era 1 had judgment but not hours; era 2 had approved words but no understanding; era 3 had understanding but no source. Era 4 is the only box that is both fluent and bound to a record.

Era 1 · live chat, staffed by people

At first, the box was a person.

The earliest chat boxes on business websites were a window into a support queue. A visitor typed; somewhere, a trained person read the message, checked what they needed to check, and wrote back. Live-chat software made this cheap enough to put on almost any site, and on plenty of sites it is still there today.

It got a surprising amount right. A person can read a messy question and work out what is actually being asked. A person knows which policy is current, can open the order, and can say the most useful words in customer service: “let me check.” When a question is unusual, a person notices that it is unusual and asks someone else. None of that is automatic. It is judgment, and judgment was the product.

What it got wrong was everything around the person. The box kept office hours, so late on a Sunday it turned grey and offered a form. It had a queue, so “you are number 6 in line” became the first thing many customers read. One agent often juggled several chats at once, which stretched every reply. Two agents could answer the same question two different ways, and a good answer, once given, lived in a single transcript where nobody could reuse it tomorrow.

What an Agentic Harness keeps from this era is the part people trust most: a real path to a person, and the honesty to use it. The harness does not pretend to be the agent who went home. It answers what it can prove, and when it can’t, it turns the question into something a person will actually pick up. That hand-off is not a failure state. It is the best idea the first era had, kept on purpose.

Era 2 · scripted decision-tree bots

Then the box got a script.

Around the middle of the 2010s, a wave of businesses put bots in the corner. Messaging platforms were opening up to automated accounts, bot-building tools were everywhere, and the pitch was simple: answer the common questions instantly, at any hour, with no queue.

Most of those bots were decision trees. A greeting, a row of buttons, a few keywords that triggered canned answers, and a form at the bottom of every branch. Every word the bot could say had been written by a person in advance and approved before it went live.

That last point is the era’s quiet virtue. A scripted bot could not invent a refund policy, because it could only say sentences somebody had written. It was predictable, reviewable, and never creative in the wrong place. If one idea from this era deserves to survive into the fourth, it is this: what a business commits to should come from a source a person approved, not from whatever sounds right.

The weakness was everything else. Real customers do not speak in buttons. Word a question a little differently from the script and the answer was “Sorry, I didn’t understand.” Ask something the tree’s authors never imagined and you went round the loop again. Every branch had to be maintained by hand, so trees went stale while the business changed underneath them. And a bot that cannot understand you often cannot tell that it should hand you to a person, either.

The frustration outlived the technology. In an October 2025 Qualtrics survey of more than 20,000 consumers in 14 countries, nearly one in five who had used AI for customer service said they got no benefit from it, and about half worried that AI would keep them from reaching a human. A box that blocks the way to a person is remembered long after a box that merely fails.

Era 3 · language-model chat with nothing around it

Then the box learned to talk, and to guess.

Large language models fixed the script’s biggest problem almost overnight. Suddenly the box could read any wording, in any tone, and reply in fluent, friendly sentences. Nobody had to write a branch for every question. For a while it looked like the end of the story.

It wasn’t, because fluency is not evidence. A model on its own produces the answer that sounds most likely, and in customer service the most likely-sounding answer is often a policy the business never wrote. Worse, a bare model will happily describe actions it has not taken: approved, refunded, shipped. With no record to check and no rule about what it may promise, the chat box became the most confident voice in the building and the least accountable one.

The public record from this era reads like a stack of error dialogs. Each case below is real and cited, and each one points at a specific piece of a harness that would have caught it.

  • The $1 Tahoe

    Chevrolet of Watsonville · December 2023

    A user told the dealership’s chatbot to agree with anything the customer said. It then agreed to sell a 2024 Tahoe for one dollar and called the offer legally binding. The bot was pulled and the “deal” was not honored.

    Control: the model never makes a commitment the system has not approved, and text typed into the box is data, never orders.

    AI Incident Database #622

  • The bot that swore

    DPD · January 2024

    After a system update, the parcel company’s chatbot swore at a customer who couldn’t get help and wrote a poem calling the company useless. DPD disabled the AI part of the bot.

    Control: tone rules, regression tests after every update, and an easy route to a person.

    ITV News, 19 Jan 2024

  • The bereavement fare

    Moffatt v. Air Canada · decided February 14, 2024

    The airline’s website chatbot told a grieving customer he could claim a bereavement fare after travelling; the airline’s own policy page said otherwise. The tribunal rejected the idea that the chatbot was responsible for itself, held the airline responsible for the information on its website, and awarded C$812.02.

    Control: answer only from the current policy source, and send exceptions to a person.

    McCarthy Tétrault analysis

  • The city’s bad advice

    NYC MyCity · reported March 29, 2024

    New York City’s business chatbot told employers and landlords things that would break city law, on topics such as tips and housing vouchers. A small disclaimer was the main safeguard.

    Control: grounding in authoritative sources, escalation on legal topics, and consistency tests that ask the same question many times.

    The Markup, 29 Mar 2024

  • The policy that didn’t exist

    Cursor support bot “Sam” · April 2025

    Asked why users were being logged out, the support bot explained the bug by inventing a one-device-per-subscription policy. Users cancelled, and a co-founder corrected it in public.

    Control: never invent policy. “I don’t know; let me get a person” beats a confident guess, and the box says plainly that it is an AI.

    AI Incident Database #1039

Notice what the fixes have in common. None of them asks for a smarter model. They ask for a source to answer from, rules about what may be promised, tests that run after every change, and an honest way out. That is the work of the system around the model, and it is where the fourth era begins.

Era 4 · the Agentic Harness chat box

Now the box answers from sources.

The fourth box keeps the third era’s way with language and puts it inside a system. That system has a name, and this library uses Deyaf’s definition of it:

“An Agentic Harness is everything around the AI model: what the assistant knows, what it remembers, where it meets customers, what it may do, and how your team teaches it.”

Walk through it on Deyaf’s site: how an Agentic Harness works

LangChain’s engineers put the same idea as a formula in March 2026: Agent = Model + Harness. In a chat box, the harness is the part that decides what the model gets to see, what it is allowed to say, and what happens next.

Deyaf describes one answer as four steps, and the model is only one of them. For a chat box, that sequence changes almost everything about the conversation you just ran:

  1. Listen. The visitor’s words arrive with context. The page they are on and the product they are looking at ride along, so “does this fit?” means something.
  2. Look it up. The exact record (the catalog entry, the manual, the warranty or the policy) is checked before any exact claim is made. A convincing product name is not evidence.
  3. Check the rules. What the assistant may say and do is decided by the system, not improvised by the model. It can explain the warranty; it cannot approve the replacement. Answering is not acting.
  4. Reply. The answer streams in as it is written, with a tile for the product or resource it came from, and one useful next step.

The fourth era also borrows back from the first two. From live chat it keeps the person: when the answer is not in the sources, the question becomes a durable item the team will see instead of a guess. From the script it keeps approval: the knowledge a business commits to is written and reviewed by people, and the harness measures whether the assistant is using it well.

Two newer expectations shape the box too. In the EU, Article 50 of the AI Act applies from 2 August 2026 and requires that people be told they are dealing with an AI system, at the latest at their first interaction; an honest box introduces itself as one, everywhere. And customers want continuity. In Zendesk’s CX Trends 2026 research (vendor-reported), 74% of consumers said they were frustrated at having to repeat information. The harness answer to that is careful memory, not endless memory, which is what the next window is about.

Get Info: what the box knows about you

Every era had a different answer to “what does the box know about me?” The person knew what you told them and whatever their screen showed. The script knew which button you pressed. A model with a long, unmanaged history can end up carrying everything, including things it should not repeat and things that are no longer true.

A harness chat box is deliberately narrow. It sees this conversation, in a bounded window, plus the page you are on. The bound is a feature: Chroma’s Context Rot study (July 2025) found that all 18 models it tested got worse as their input grew longer, so stuffing every past message into the prompt makes answers less reliable, not more.

If you are a returning customer, short historical notes can help it pick up the thread, but notes are never treated as current facts. An order’s status, a shipment or a balance is looked up fresh from the live record every time. Contact details are stored apart from published knowledge and never copied into it.

Try to tick the empty boxes in the dialog. It will tell you why each one stays empty.The dialog lists what stays unticked, and why.

Get Info dialog: what a harness chat box knows about a website visitor. A design illustration.
Website visitorseen by a harness chat box
Kind:
One conversation
Where:
The page and product they are viewing
History:
This conversation, bounded
Notes:
Short and historical, never current facts
Orders:
Looked up fresh, every time
Contact:
Kept apart from published knowledge

Bounded by design: the box keeps this conversation in a limited window, because longer input makes answers worse.

Notes are history, not facts: order status, shipments and balances are looked up fresh from the live record.

The conversation record is for the business’s operators and is not searchable by visitors.

Contact details are stored apart from published knowledge and never copied into it.

Never: the box says it is an AI assistant, and in the EU the law requires it.

Locked. Tick a box to see why it stays empty.

The Trash that never empties

Chat boxes get things wrong, and so do the people who teach them. A lesson that was right in March can be wrong by September because the policy changed. The tempting fix is to delete it. A harness does something more careful: corrections, not deletions. The wrong lesson is archived so it stops shaping answers, a corrected lesson replaces it once a person approves it, and the record of what was said, and why it changed, stays put.

That matters more than it sounds. If a customer comes back and says “your chat box told me 60 days,” someone needs to be able to see what the library said at the time, and when it changed. A box that deletes its own history cannot be checked. A box that archives can.

Illustrative example. Drag a lesson onto the Trash, or select it and press “Move to Trash”. Then try “Empty Trash”.Three taught lessons for a fictional store; the first one is out of date.

  • Nothing is ever emptied here. A trashed lesson is archived: it stops being used in answers, and it stays on record.

Empty. Trashed lessons land here, archived.

Select a lesson to begin.

In EVE’s case a taught lesson can be quarantined and restored, and answering a customer and creating a reusable lesson are separate operations. Nothing a visitor types becomes public knowledge on its own; people write the library, and the harness measures whether it is being used well.

The working example

About this EVE

EVE is Feniex’s assistant, and the clearest way to see a fourth-era chat box is to look at one that is working. EVE runs on the Agentic Harness Deyaf packages: her knowledge, memory, doors, rules and learning loop. You can meet EVE, Feniex’s assistant on Deyaf’s site.

On Feniex’s website, EVE’s chat box streams its answer as it is written, knows which page and product the visitor is looking at, shows product and resource tiles, and keeps a bounded conversation history. There is a website voice too. Behind every door (website chat, website voice, phone, text and email) it is the same EVE, with the same identity, the same library and the same rules.

Her manner is the history on this page turned into habits. She answers first, gives one useful next step, then stops; replies are usually two or three sentences. She asks for one missing detail at a time. Before any exact product, compatibility, warranty, price or software claim she checks the applicable record, and she says only what the result proves: saved is not sent, and sent is not resolved. She never pretends to be a person.

When she doesn’t know, the question doesn’t vanish. An unanswered question becomes a durable item for the team; when the visitor has left contact details, a support ticket is filed, and she says it is “with the support team” only once the ticket is confirmed accepted.

EVE · Feniex’s assistantDoors: website chat · website voice · phone · text · email
  • Chat: first word1.4 s
  • Chat: finished reply2.7 s
  • Handled well (100 fixed test questions)86%
  • Material errors, same test10%

As of September 24, 2026 · Feniex's internal Eve 3.0 report · machine-judged and provisional. Speed bars are scaled to 3 seconds; report card graded September 23, 2026. EVE’s report card, weak spots included

Read those numbers as a pair. On 100 fixed test questions graded September 23, 2026, 86% were handled well and 10% contained a material error; the second figure is the reason the first is not the whole story. In speed terms, that is about 3 seconds to a finished reply in website chat and on the phone, which Deyaf’s homepage walks through as four steps in about three seconds.

The same harness answers Feniex’s phone when the team is busy or the office is closed. In her first week on the line, since September 17, 2026, EVE answered 102 calls: 86 while the team was busy and 16 after the office closed. 93% of callers stayed and talked, and 60 were handed to the team with the caller’s details taken; that is a hand-off, never a transfer, because EVE does not transfer calls by design (Feniex’s own measurements, Sep 17–24, 2026; see the dated report card). A chat box and a phone line look like different products. Under a harness, they are two doors into one assistant.

Help topics

Choose a topic to open its “About…” box.

About chatbots and harnesses…Is an Agentic Harness chat box just a better chatbot?

From the visitor’s side it looks like one: the same box, the same typing. The difference is the system behind it. A chatbot is usually a script or a bare model. A harness chat box is a model inside something that supplies sources, enforces rules, keeps memory bounded and routes what it can’t answer to people. On Deyaf’s site, the model is one step of four.

About going back to scripts…Scripts never invent anything. Why not just use a decision tree?

For a handful of fixed tasks, a script is still fine. It breaks on real wording and on questions nobody predicted, and it tends to trap people in loops. The harness keeps the script’s best property, answers that come from approved sources, and drops its worst, the dead end.

About memory…Will the chat box remember me next time?

Within a conversation, yes, inside a bounded window. Across visits, a harness keeps short, historical notes at most, and it looks up anything current (an order, a shipment, a balance) fresh from the live record. The full conversation record is for the business’s operators, not for other visitors. Deyaf lays out four kinds of memory, each with one job.

About saying it’s AI…Does the box have to say it’s an AI?

In the EU, Article 50 of the AI Act says people must be told they are interacting with an AI system, and it applies from 2 August 2026. Outside the EU it is still the honest choice: the Cursor case shows what happens when a support bot sounds like staff and invents policy. EVE never pretends to be a person.

About not knowing…What happens when it doesn’t know?

It says so. The unanswered question becomes a durable item the team will see, and with the visitor’s contact details a ticket is filed. Confirmation is only given once the ticket is accepted. Deyaf’s homepage shows what happens when she doesn’t know.

About building one…Can I build one of these for my own website?

Deyaf’s current release is an early-access setup and preview experience. The builder has four steps (your business, what it knows, how it helps, try it) and an exact-source knowledge preview that runs in your browser; it is not live AI. Trying it takes about ten minutes with no account, starting from a front desk that knows. Live doors are switched on one at a time after activation, and Deyaf explains how to switch on each door deliberately.

Build the fourth box?

Every era of the chat box left something worth keeping: the person, the approved words, the fluent reading. An Agentic Harness keeps all three and answers from your sources. Deyaf’s builder lets you set one up and try it in your browser.