Deflection rate is the most quoted and least useful number in AI support. This report separates it into components that mean something operationally: conversations contained and genuinely resolved, conversations contained and abandoned, and conversations handed to a human with or without loss of context.

What we measured

The dataset covers automated conversation outcomes across a broad sample of OMNELIAS accounts, classified by intent family and followed for a fourteen-day window to detect follow-up contacts. That follow-up window is the critical methodological choice: without it, an abandoned conversation is indistinguishable from a resolved one.

Findings that change roadmaps

  • Resolution quality varies far more by intent family than by model. Account and order status intents behave completely differently from billing disputes, and averaging them hides the difference.
  • Hand-off satisfaction is the strongest predictor of whether customers accept automation at all, and it depends almost entirely on whether context survives the transfer.
  • Confidence thresholds set too low produce a characteristic signature: high containment, high abandonment, flat satisfaction. It is visible within a fortnight if you measure for it.
  • Teams with formal quality assurance on AI-drafted replies expanded automation scope faster than teams without, because they could evidence safety.

Using the benchmarks

Each intent family has a table with the observed distribution of contained-and-resolved rates, so you can compare your own performance per intent rather than in aggregate. The final section is a diagnostic: given a pattern of numbers, which of five common failure modes you are most likely experiencing, and what to change first.