Frederik Rybansky

AI InfrastructureAI AgentsBratislava, SK

What an AI Chatbot on a Website Actually Returns

Chatbots fail commercially in one specific way: nobody measured the right thing, so the pilot survived on conversation counts while support volume stayed exactly where it was.

A website assistant is the easiest AI project to demo and the hardest to justify. Both facts come from the same cause: everyone measures the easy number — how many people used it — instead of the hard one — how much work it removed.

Why most website chatbots get cancelled

The typical lifecycle looks like this. Someone builds a widget in a few weeks. Traffic is high, so conversation counts look impressive. Support volume does not move. Twelve months later the budget is cut and the organisation concludes that AI does not work.

What actually happened is that the assistant was answering the questions it was good at, and the questions it was good at were not the expensive ones. Visitors ask what they can find in the footer. Those questions are real, and answering them is pleasant, and they were never costing you anything.

The expensive questions — refund eligibility, warranty scope, whether a plan covers their specific case, what happens to their data — are exactly the ones where a wrong answer creates a liability. Those are the questions the assistant was configured to refuse.

Deflection: the number nobody measures properly

Deflection is the fraction of a contact channel that never happens. It sounds simple and it is the number that decides the business case. Measuring it properly requires three things:

A baseline. Support tickets and chat volume per week for at least four weeks before launch, split by question type. Without this you have a feeling, not a number.

A definition of counterfactual. When a visitor uses the assistant and then emails support anyway, the contact did not happen — or it happened faster and cheaper because the agent arrived with context. You have to decide which, in writing, before launch, and apply it consistently.

A holdout. A small share of visitors who see the assistant and a small share who do not, both measured. This is the only method that survives contact with seasonality, because it controls for a slow week automatically.

Without a holdout, every external event — a marketing push, an outage, a price change — becomes an explanation you reach for when the numbers disappoint.

Resolution versus deflection

These are different and only one of them is honest.

Deflection is "the visitor stopped and did not contact us". Resolution is "the visitor got what they needed and would say so". Deflection can be achieved by a confusing assistant that frustrates people into emailing anyway, which is why deflection without a satisfaction or follow-up signal measures nothing useful.

Track both. A high deflection rate with a low resolution rate is a warning, not a success. The simplest useful signal is whether visitors contact you about the same topic less afterwards, which the holdout makes measurable.

Instrument it from day one

Before you build anything, decide the events. A reasonable minimum:

  • Widget opened, conversation started, conversation abandoned.
  • Question asked, resolved from cited content, or escalated to a human.
  • Language used, page the visitor was on, topic detected.
  • Post-conversation outcome: contact made or not made within seven days.
  • An explicit thumbs up or down on each substantive answer.

Anonymous, first-party, no third-party trackers. You do not need a tracking vendor to answer these questions, and adding one creates a consent problem you would then have to solve.

The unanswered-question list

The most valuable artefact a website assistant produces is not the answers — it is the list of questions it could not answer with confidence.

Publish it internally, weekly. It is simultaneously a content backlog, a product roadmap, and the business case for the next phase. In my experience a third of the first quarter's list gets fixed by editing documentation, which costs hours and is the cheapest possible improvement.

That list is also the honest defence of the project. "It answers 78% of the questions it sees, and here are the twenty questions it cannot, and here is who is fixing them" is a conversation an executive committee can have. "It got 40,000 chats" is not.

Multilingual changes the maths

Adding Czech and Slovak roughly doubles your addressable question set, and it changes its shape. Local visitors ask about local pricing, local warranty law, and local returns — content many English-first sites simply do not have.

This produces a distinct pattern: the assistant performs well on English, and on Czech it has a lower answer rate for reasons that are about your documentation rather than about the model. Instrument per language, always. A single blended metric hides exactly the problem you most need to see.

It also means the multilingual rollout is frequently justified on content grounds alone, independent of the assistant — the questions coming out of the Czech logs are a specification for your Czech website.

What a good answer looks like

Three properties, in order of importance:

  1. It cites. A source the visitor can open and check. This converts "trust me" into "verify this".
  2. It admits ignorance. When the corpus does not answer, say so and offer the contact route. A refusal is a correct answer; an invention is a defect.
  3. It is short. Your support articles are written for browsing, not for reading inside a chat bubble.

Everything else — tone, personality, speed — is worth far less than those three, and much easier to fix.

When to turn it off

Turn it off, or narrow it, if any of these are true after a full quarter with a holdout:

  • Support volume on covered topics has not moved.
  • Escalation-to-human rate is climbing rather than falling.
  • Answer quality scores are flat while conversation volume grows — that combination means people are asking harder questions and it is guessing.
  • Maintenance cost of keeping the content corpus accurate exceeds the cost of the conversations it saves.

That last one is real. A website assistant requires someone to keep documentation current. If nobody owns that, the system degrades quietly, and a degraded assistant that still answers confidently is worse than no assistant at all.

Related writing