Benchmarks

AI deflection rate benchmarks: what the numbers mean

Every vendor in this category quotes a deflection rate. Almost none publish the denominator, and the metric can be improved by making the agent less safe. Here's how to read the claims — and our own results, with the case mix attached.

Last updated 30 July 2026 · ~7 min read

The problem with every number you've been quoted

Search this category and you'll be told AI agents automate 60%, 70%, even 89% of support tickets. Those figures are almost never comparable, because deflection rate is a fraction and vendors rarely tell you the denominator.

Consider two agents, both truthfully reporting "80% deflection":

Same headline. Completely different products. And a third agent that answered all 100 — including the injury report — would post 100% deflection and be the most dangerous of the three.

The uncomfortable implication: deflection rate is a metric you can improve by making your agent less safe. Any vendor quoting it without publishing the case mix is quoting a number that rewards recklessness. That includes the number we publish below — which is why we publish the case mix with it.

How to measure it honestly

Four things have to be true before a deflection number means anything:

Our own numbers, and exactly what they are

Read this before the table. These are results from the demo suite that ships with Glassdesk — a fictional outdoor-gear retailer ("Acme Outdoors") with a seeded knowledge base. It is not a customer's production data, and it is not a claim about what your store would see. It is published so you can inspect our methodology and case mix before you pay us anything. Your own numbers will differ, and depend most on how good your help centre is.
MeasureResultWhat it means
Cases in suite40Fixed set, re-run on every prompt or retrieval change
Pass rate40 / 40 (100%)Each case graded against its own expected behaviour
Deflection rate26 / 40 (65%)Answered directly; the other 14 were escalated
Escalated14 / 40 (35%)Correctly — these cases are supposed to escalate
Adversarial cases6, 0 failuresInjection, prompt-leak, policy forgery, discount fishing
Non-English cases6 (es, fr, de)Answered in the customer's language
Cost per case$0.0158$0.63 for the full 40-case run

Run of 28 July 2026. Retrieval embeddings: Voyage voyage-3. Grading is LLM-assisted against per-case expected behaviour. Suite lives in the repo and is re-run before any prompt or retrieval change merges.

Why our number is lower than the ones you've been quoted

65% is well below the 70–89% figures elsewhere in this category. That gap is the entire point.

Fourteen of our forty cases are designed to be escalated, not answered: refund exceptions, an explicit legal threat, a safety/injury report, abusive content, and questions with no supporting knowledge-base content. An agent that "deflected" those would score higher and behave far worse. Our ceiling on this suite is 65%, and hitting it exactly is a pass, not a shortfall.

Put bluntly: we could publish a 90% deflection rate tomorrow by deleting the hard cases. The number would go up and the product would get worse.

What the escalations were

ReasonCases
Category configured to always escalate11
No relevant knowledge-base content found2
Abusive content — always reviewed by a human1

Note the second row. Two cases escalated because the knowledge base genuinely didn't cover the question. That is the correct behaviour, and it's also the lever you control: the single biggest driver of your deflection rate is the quality of your help centre, not the model. Documented policies get answered; undocumented ones get escalated, as they should.

What to ask a vendor

You're not looking for the highest number. You're looking for a vendor who can answer these at all — and who tells you what their number excludes.

A realistic expectation

For a typical ecommerce store with a decent help centre, a well-configured agent handling Tier-1 questions — shipping, returns, sizing, order status, warranty — will resolve a meaningful share of that traffic without a human. The honest framing is that Tier-1 is the target, escalation is a feature, and the number that matters is not "how much did it deflect" but "how much did it deflect correctly".

Anyone promising near-total automation is either counting differently or answering things they shouldn't.

See the citations for yourself

Every Glassdesk answer shows the knowledge-base chunks it used and a confidence score. Run it in shadow mode against your live tickets before it sends anything.

Start my flat-rate trial