How-to guide

Shadow mode: test an AI support agent without risking a customer

The real objection to AI support isn't cost. It's that a bad answer goes out under your brand name, to a customer you spent money to acquire, and you find out afterwards. Shadow mode removes that risk entirely — the AI answers your real tickets, and only your team can see what it said.

Last updated 13 August 2026 · ~5 min read

What shadow mode actually does

The AI runs the full job on your genuine, live tickets: it reads the customer's question, searches your help centre, writes an answer, and scores its own confidence. Then, instead of sending, it posts that answer as an internal private note on the ticket.

Your agent opens the ticket, sees the draft sitting there, and either uses it or ignores it. The customer sees nothing at all. A wrong answer costs you the fifteen seconds it took to read it.

Why this beats a sandbox. A demo or test environment runs on tickets somebody made up — and invented tickets are always tidier than real ones. Real customers write in fragments, ask three things at once, attach a photo instead of explaining, and reply "still not working" four days later. Shadow mode is the only test that uses those. Everything is real except the send.

The four things to measure

Run it for one to two weeks — long enough to catch your ordinary questions, at least one busy day, and a few genuinely awkward tickets. Then judge it on these:

MeasureThe question it answersWhat good looks like
Send-unchanged rateHow often would an agent have sent the draft as written?The number that decides whether this saves you time
Escalation rateHow often did it refuse and hand off instead of guessing?Should be non-zero — see below
Citation qualityDo the linked articles actually support the answer?Every claim traceable to a real article
Confidently wrongWas it ever certain and incorrect?Zero. This is the only hard blocker

A high escalation rate is not a failure

This surprises people, so it's worth being direct about it. If the AI escalates a ticket to a human because your help centre doesn't cover the question, the AI is working correctly and your documentation has the gap. An agent that answers everything is an agent that is inventing things.

Read the escalations as a list. They are the most useful output of the whole trial: a ranked, evidence-based list of the articles your help centre is missing, written by your actual customers. Most teams find the two-week escalation log more valuable than the deflection number.

What to do at the end of two weeks

Questions to ask any vendor about this

How Glassdesk does it

Shadow mode is the default here, not an option — every new connection is created with drafts on and sending off, for Gorgias and Freshdesk alike. You switch it off deliberately, once you've read enough drafts to trust it.

Alongside it, every answer carries the help-centre articles it used and a confidence score, so a draft you disagree with can be traced to the article that caused it — which is usually a documentation fix rather than a model problem. Below your confidence threshold it escalates rather than guessing, and any retrieval or tool error escalates too, so a customer is never left on a dead end.

Pricing is flat — $79 / $199 / $449 — so a shadow-mode week costs the same as a live one, and reading drafts is never billed as work.

Watch it work before it talks to anyone

Connect your helpdesk, let it draft against your real tickets for two weeks, and read every answer with its sources. Nothing reaches a customer until you say so.

Start my flat-rate trial