Support quality assurance was designed for a world where a human wrote every word. Sample five conversations per agent per month, score them against a rubric, coach the gaps. When an AI drafts the majority of replies and an agent edits and sends them, that process measures the wrong thing: you end up grading a human for a model's phrasing, and grading nobody for the model's judgement.
Score the reply, then attribute it
Keep one rubric for the customer-facing outcome, and record separately who produced the text. A workable five-dimension rubric:
- Accuracy: is every factual claim supported by product documentation or account data?
- Completeness: does it resolve the question asked, including the implied one?
- Policy: does it stay inside commercial and data-handling rules?
- Tone: does it match the customer's register and the severity of their situation?
- Efficiency: could the same outcome have been reached in fewer messages?
Then tag each scored reply as model-authored, model-drafted-and-edited, or human-authored. The interesting signal is the delta between those three groups on each dimension, which tells you precisely where the model needs work and where your agents need coaching.
Sample by risk, not at random
Random sampling wastes reviewer time on the easy majority. Weight your sample towards conversations that carry risk: refunds and credits, policy exceptions, low sentiment, reopened threads, unusually short handle times, and any conversation where the agent sent an AI draft without editing a single character. That last cohort is worth watching closely, unedited sends are where automation bias shows up first.
Grade the hand-offs too
The transition from AI to human is a distinct artefact and deserves its own score. Did the summary carry the actual problem? Did the customer have to repeat anything? Did the human contradict what the AI already promised? Contradiction is the most damaging of the three and the easiest to catch in review.
Close the loop into the knowledge base
Every accuracy failure should produce one of three outputs: a documentation fix, a retrieval fix, or a guardrail. If a reviewer marks something inaccurate and nothing changes upstream, the same error will recur at scale next week. Route QA findings into the same backlog that owns your help centre content, and track the median time from finding to fix.
A cadence that works
- Weekly: a risk-weighted sample of 30 to 50 conversations, reviewed by rotating senior agents.
- Monthly: a calibration session where reviewers grade the same five conversations and reconcile differences.
- Quarterly: a full re-grade of the previous quarter's worst dimension to confirm the fixes held.
Done consistently, QA stops being a compliance ritual and becomes the mechanism that lets you safely widen the scope of automation.