Actually Helpful

Why '99% accurate' chatbot claims can't be taken at face value

A high accuracy percentage sounds precise, but usually hides unstated assumptions about what "accurate" means and how it was measured.

Published August 19, 2026

The number that's supposed to settle the question

Somewhere on a chatbot vendor's pricing page or sales deck, there's usually a single figure meant to end the debate before it starts: "99% accurate." It reads like a fact you can just take at face value and move on.

Ask what was actually measured, though, and the claim underneath the number tends to shrink. An answer that matched the vendor's own reference material, one that didn't happen to contradict a known fact, or one a reviewer skimmed and didn't flag as obviously wrong can all get rounded up into the same headline figure.

None of those is the same as what you assumed you were buying: a correct, useful answer to your actual question. That gap matters when you're deciding whether to trust a tool with your support traffic.

What "accurate" usually actually means in these numbers

Start with the definition of "accurate" itself. Vendors often define accuracy loosely, and the definition shapes the entire number.

Does "accurate" mean the answer is perfect? Probably not. Often it means "not obviously wrong" or "contains some correct information" or even just "the chatbot didn't refuse to answer." You can score very high on almost any metric if you're generous about what counts.

Does it mean the answer is complete and usefully specific to your question? Usually not. A vague answer that happens to be technically true can count as accurate.

Does it mean a human customer actually found the answer helpful? Rarely. The vendor tested how the answer looks on paper, not whether a real person with a real problem found it solved anything.

The definition of accuracy is doing most of the real work behind the claim. A 99% number sounds like a guarantee of quality. The actual definition usually guarantees much less.

Who decided what questions to test

Here's the second hidden assumption: the set of questions the vendor tested against.

Many vendors test against their own curated set of questions. These are often the common, straightforward questions they know the chatbot handles well. They're picked to make the vendor's tool look good.

Real customer support questions are messier. They're vague. They ask about situations the docs don't quite cover. They contain typos and abbreviations. They assume context the chatbot doesn't have. Some ask for things outside the chatbot's scope on purpose, testing what it does when stuck.

A chatbot that scores 99% on clean, curated questions from its vendor might score lower on a random sample of real questions from real customers. The difference between those two numbers is not a measurement failure. It's a gap between what the vendor chose to measure and what actually matters to you.

Ask a vendor directly, and the exchange tends to go something like this:

You: Your documentation says your chatbot is 99% accurate. What's that number based on?

Vendor rep: We test it against 50 sample support questions from our own knowledge base, checking the answers are on topic and don't contain obvious errors.

You: What if I tested it against 50 of my own support questions, including the messy ones?

Vendor rep: We can't guarantee it would score the same. We know it performs well on typical questions related to our product.

That exchange is honest, even if it's awkward for the vendor. It explains what the 99% actually measures. Not every vendor offers that level of clarity upfront, which is itself useful information when you're evaluating a tool.

Other assumptions hidden in the number

An accuracy percentage usually comes with several other unstated decisions that shape it:

Questions worth asking a vendor who cites an accuracy number

  1. What specific definition of "accurate" are you using? Does the answer have to be complete and correct, or just not obviously wrong? What was the actual scoring rule?
  2. Who picked the test questions? Were they chosen by your team, or drawn from a random sample of real support tickets? Would you test against my hardest questions instead?
  3. How many questions were tested? What's the actual sample size, and how does that compare to the volume of questions you'd realistically send it in a month?
  4. Who scored the answers? A person, a machine, or a mix? Were multiple reviewers involved, and how consistent were they with each other?
  5. Has this been checked by anyone outside your team? Independent verification adds real weight to a claim like this.
  6. Would it score the same on a different, harder set of questions? This is the test that matters most for your actual use case.

Where we land on this

We don't publish a headline accuracy percentage as a marketing claim. Here's why, and what we do instead.

We run our own test suite before we treat any engagement as ready: about 65 test questions generated from your own documentation, layered with deliberately vague, out-of-scope, and hard ones we wrote to try to break it. We check each answer against known failure patterns, not just whether it looks right on a quick read.

Whenever we do report a measured number to anyone, we state what it's based on: the actual sample size and what was tested. A number with no stated basis isn't a claim worth making, and we try not to make it.

We've also made a standing offer: send us your own hardest real support questions and we'll run them through and send back the answers unedited, including any we get wrong. You see the real output, and you can check it against your own docs yourself instead of trusting a percentage.

A single figure like that is easy to print on a slide and hard for anyone else to check. A real answer with a real source attached is the opposite: less punchy, but you can confirm it yourself instead of taking our word, or anyone's, for a number.

Related reading

Test it yourself instead of taking a number's word for it

A percentage is easy to publish and hard for you to double check. A real question and a real answer aren't. Bring us one and see what actually comes back.