Quality scores are only useful if they mean the same thing regardless of who assigned them. Most support teams discover the hard way that their reviewers disagree substantially on identical conversations, often by a full point on a five-point scale. Calibration is the cheap, unglamorous practice that fixes this, and it takes an hour a month.
The format
- Before the session: the QA owner picks five conversations, two straightforward, two genuinely contested, one involving an AI-drafted reply. Every reviewer scores them independently and submits before seeing anyone else's scores.
- First 10 minutes: reveal the score spread per dimension. Do not discuss reasoning yet. Simply seeing the variance sets the tone.
- Next 35 minutes: work through the two or three widest disagreements. The reviewer at each extreme explains their reading; the group agrees a canonical interpretation.
- Final 15 minutes: write the outcome down as rubric amendments with concrete examples. If nothing gets written, the session did not happen.
Facilitation rules that matter
Score before you discuss, always, anchoring destroys the signal otherwise. Anonymise the agent whose conversation is being reviewed, so the discussion is about the reply rather than the person. Timebox each conversation strictly; the ninth minute of debate about tone has never produced a rubric improvement. And rotate the facilitator, because a single owner's interpretation gradually becomes the standard by default rather than by agreement.
Track whether it is working
Measure inter-rater agreement on the calibration set each month. A team starting cold typically sits well below acceptable agreement and reaches a workable level after three or four sessions. Publish the trend. It is the evidence that your quality scores can be used for coaching and, eventually, for decisions about automation scope.
Extend it to AI output
Include at least one AI-drafted reply in every session. Reviewers apply consistently harsher standards to model output at first, and the discussion about why is one of the most productive conversations a support team can have. It is also how you arrive at a defensible answer to the question your executives will eventually ask: how do we know the automated replies are good enough?