Actually Helpful

What 'deflection rate' numbers actually measure (and don't)

You've probably seen a support AI vendor tout a specific percentage for how many conversations their chatbot 'deflects' from human agents. Here's what that number actually tells you, and what it doesn't.

A number that sounds more solid than it is

You've probably seen a chatbot vendor advertise something like "deflects most support tickets" or a specific deflection percentage. It sounds like a hard, measured fact. It usually isn't, or at least, it's measuring something much narrower than it sounds like.

The gap between what "deflection rate" sounds like it means and what it actually measures is where a lot of vendor claims get their power. It sounds like the chatbot solved most of your support problem. What it often actually means is something much smaller: that a visitor didn't click a button to talk to a human, or didn't ask for an agent in that particular conversation.

That's not nothing, but it's not the same thing, and the difference matters when you're deciding whether to invest in a tool.

How a deflection rate usually actually gets calculated

Here's the plainest explanation: a common way vendors calculate deflection is simply counting conversations that ended without the visitor clicking "contact support" or explicitly asking for a human. That's it.

It does NOT mean the visitor got a correct answer. It does NOT mean they got a helpful answer. It does NOT mean they were satisfied or that their problem was actually solved. It just means they didn't take one specific next action in that moment.

Think about what can look like a "deflected" conversation under this definition. Someone who got a wrong answer and gave up counts. Someone who got frustrated and left the site entirely counts. Someone who got a genuinely great answer and solved their problem counts. Someone who decided they'd rather not deal with support at all today counts.

All of those conversations "succeeded" on the metric of ending without an escalation. But the outcomes are completely different. One visitor actually got helped. Three got nothing useful. The metric treats them the same.

It's like measuring the success of a restaurant by counting how many people didn't complain to the manager, without ever checking whether the food tasted good. Some people just don't complain. That doesn't mean the kitchen is working.

What a real, trustworthy number would require

To actually know if a chatbot is helping (not just ending conversations), you'd need something much more rigorous. You'd need a control group, some customers who use the chatbot and a truly comparable group without it, and a way to measure real outcomes. Did they actually get their problem solved? Did they come back with the same question later? Were they satisfied with the interaction?

You'd need to follow up with users and ask them whether they felt helped. You'd need to track whether they returned to support for the same issue. You'd need to compare ticket volume and resolution rates across a real baseline and your experimental group, accounting for the dozen other variables that change month to month.

Controlled, methodology-disclosed comparisons like that are rare to find published anywhere, and running one properly is genuinely hard and expensive. It requires discipline, time, and a willingness to publish a smaller number if the honest answer is smaller. Most businesses understandably just measure what they can measure easily.

Why this number still gets used everywhere

This isn't necessarily dishonesty on purpose. Deflection is just genuinely easy to measure (count the conversations that didn't escalate) compared to the real thing you actually care about (did this genuinely help this person).

Easy-to-measure numbers tend to get used even when they're a weak stand-in for the thing that actually matters. It's a pretty universal pattern. You measure what's simple to count, it gets repeated, it becomes the standard everyone cites, and pretty soon it's the metric that defines the category. Everyone gets used to seeing "deflection rate" numbers because they're cheap to calculate and presentable on a slide.

That doesn't make them less real as a number. It just means they're measuring something smaller than the name suggests.

Questions worth asking any vendor who cites a deflection number

  1. How exactly is "deflected" defined here? What specific action (or lack thereof) counts as deflection? Did the visitor have to avoid clicking a button, or is simply not asking for a human enough?
  2. Was this measured against a real control group? Did you compare a period with the chatbot to a truly equivalent period without it, or just look at before and after numbers on the same system?
  3. Does this account for people who just gave up? Is there any way to distinguish between someone who got a great answer and someone who got frustrated and left?
  4. Can they show you real examples, not just the aggregate number? Ask them to pick a few actual conversations and walk you through them. Did the visitor actually get what they needed?

Where we land on this

We don't publish a deflection rate or an ROI number for our own product. We think a number like that would overstate what we can honestly claim, especially without a control group and without following up to verify that people actually got what they needed.

What we can show you instead: a real question, answered from your real docs, with a real source you can check yourself. That's slower to look at than one big percentage. It doesn't fit neatly on a slide. But it's actually true, and you can verify it yourself instead of just trusting a number.

When you're evaluating a support tool, that's a better foundation than a percentage you can't reproduce and don't fully understand.

See a real answer instead of a big percentage

We'd rather show you one real, checkable answer than one big number you can't verify. Ask our agent a real question and judge for yourself.