The number that's supposed to settle the question
Somewhere on a chatbot vendor's pricing page or sales deck, there's usually a single figure meant to end the debate before it starts: "99% accurate." It reads like a fact you can just take at face value and move on.
Ask what was actually measured, though, and the claim underneath the number tends to shrink. An answer that matched the vendor's own reference material, one that didn't happen to contradict a known fact, or one a reviewer skimmed and didn't flag as obviously wrong can all get rounded up into the same headline figure.
None of those is the same as what you assumed you were buying: a correct, useful answer to your actual question. That gap matters when you're deciding whether to trust a tool with your support traffic.
What "accurate" usually actually means in these numbers
Start with the definition of "accurate" itself. Vendors often define accuracy loosely, and the definition shapes the entire number.
Does "accurate" mean the answer is perfect? Probably not. Often it means "not obviously wrong" or "contains some correct information" or even just "the chatbot didn't refuse to answer." You can score very high on almost any metric if you're generous about what counts.
Does it mean the answer is complete and usefully specific to your question? Usually not. A vague answer that happens to be technically true can count as accurate.
Does it mean a human customer actually found the answer helpful? Rarely. The vendor tested how the answer looks on paper, not whether a real person with a real problem found it solved anything.
The definition of accuracy is doing most of the real work behind the claim. A 99% number sounds like a guarantee of quality. The actual definition usually guarantees much less.
Who decided what questions to test
Here's the second hidden assumption: the set of questions the vendor tested against.
Many vendors test against their own curated set of questions. These are often the common, straightforward questions they know the chatbot handles well. They're picked to make the vendor's tool look good.
Real customer support questions are messier. They're vague. They ask about situations the docs don't quite cover. They contain typos and abbreviations. They assume context the chatbot doesn't have. Some ask for things outside the chatbot's scope on purpose, testing what it does when stuck.
A chatbot that scores 99% on clean, curated questions from its vendor might score lower on a random sample of real questions from real customers. The difference between those two numbers is not a measurement failure. It's a gap between what the vendor chose to measure and what actually matters to you.
Ask a vendor directly, and the exchange tends to go something like this:
You: Your documentation says your chatbot is 99% accurate. What's that number based on?
Vendor rep: We test it against 50 sample support questions from our own knowledge base, checking the answers are on topic and don't contain obvious errors.
You: What if I tested it against 50 of my own support questions, including the messy ones?
Vendor rep: We can't guarantee it would score the same. We know it performs well on typical questions related to our product.
That exchange is honest, even if it's awkward for the vendor. It explains what the 99% actually measures. Not every vendor offers that level of clarity upfront, which is itself useful information when you're evaluating a tool.
Other assumptions hidden in the number
An accuracy percentage usually comes with several other unstated decisions that shape it:
- Sample size. Was the test 10 questions? 100? A tiny sample can produce a high number just by chance, and a 99% figure from a handful of questions means something different from a 99% figure measured across a much larger set.
- Who checked it. Did a human reviewer look at each answer? Did they use a consistent checklist, or was it a quick visual scan? Human judgment is real and honest, but it's also variable, and that variability shapes the number.
- What it was tested on. Was it tested on current docs, or docs from months ago? Was it tested on the exact version of the tool you'd actually be using, or a demo version? The further the test is from your reality, the less the number predicts your experience.
- Whether it was independently verified. Did the vendor measure its own tool, which introduces an incentive to be generous with scoring? Did anyone outside the vendor check the work? The absence of independent verification says something worth noting about the claim.
Questions worth asking a vendor who cites an accuracy number
- What specific definition of "accurate" are you using? Does the answer have to be complete and correct, or just not obviously wrong? What was the actual scoring rule?
- Who picked the test questions? Were they chosen by your team, or drawn from a random sample of real support tickets? Would you test against my hardest questions instead?
- How many questions were tested? What's the actual sample size, and how does that compare to the volume of questions you'd realistically send it in a month?
- Who scored the answers? A person, a machine, or a mix? Were multiple reviewers involved, and how consistent were they with each other?
- Has this been checked by anyone outside your team? Independent verification adds real weight to a claim like this.
- Would it score the same on a different, harder set of questions? This is the test that matters most for your actual use case.
Where we land on this
We don't publish a headline accuracy percentage as a marketing claim. Here's why, and what we do instead.
We run our own test suite before we treat any engagement as ready: about 65 test questions generated from your own documentation, layered with deliberately vague, out-of-scope, and hard ones we wrote to try to break it. We check each answer against known failure patterns, not just whether it looks right on a quick read.
Whenever we do report a measured number to anyone, we state what it's based on: the actual sample size and what was tested. A number with no stated basis isn't a claim worth making, and we try not to make it.
We've also made a standing offer: send us your own hardest real support questions and we'll run them through and send back the answers unedited, including any we get wrong. You see the real output, and you can check it against your own docs yourself instead of trusting a percentage.
A single figure like that is easy to print on a slide and hard for anyone else to check. A real answer with a real source attached is the opposite: less punchy, but you can confirm it yourself instead of taking our word, or anyone's, for a number.
Related reading
- See all 20 guides in the Actually Helpful library, grouped by what you are actually trying to figure out.
- What "deflection rate" numbers actually measure (and don't), another case where a percentage hides what it's actually counting.
- Why chatbot reviews look so different depending on where you check, on how the source behind a number shapes what it means.
- How many of your support tickets are actually repetitive, for a number worth measuring yourself instead of taking a vendor's word for it.
Test it yourself instead of taking a number's word for it
A percentage is easy to publish and hard for you to double check. A real question and a real answer aren't. Bring us one and see what actually comes back.