Results
- 70%
- Tickets auto-resolved
- 4h to 2min
- First response time
- 4.2 to 4.5
- CSAT
Sustained over six months
Measured post-resolution
The challenge
Sixty percent of the inbound queue was password resets, billing questions and configuration lookups. Those tickets consumed most of the team's capacity, so genuinely complex issues sat for days.
How we approached it
Classified six months of history
We labelled the historical queue to identify which intents were both high-volume and reliably automatable before writing a line of code.
Built a cited retrieval pipeline
Retrieval runs over product documentation and internal runbooks, with a citation required on every response the agent returns.
Set an escalation threshold
Below a confidence floor the system hands off to a human with the full conversation and its retrieved context attached.
Gated deploys on evaluation
An evaluation suite of 180 labelled tickets runs on every prompt change and blocks deployment on regression.
The outcome
Seventy percent of inbound tickets now resolve without human involvement. Median first response fell from four hours to two minutes, and CSAT rose because complex tickets finally got attention.
What the client said
“We expected deflection. What we did not expect was our satisfaction score going up at the same time.”
What we would do differently
Our initial confidence threshold was too permissive. The first two weeks produced a handful of confidently wrong billing answers. Raising the threshold cut automation from 78% to 70% and eliminated the failure mode.
Frequently asked



