Quality programmes fail for boring reasons: the rubric is ambiguous, the sample is random, and reviewers never reconcile their scores. This pack is the working set of documents we recommend to fix all three, with the wording deliberately concrete enough to use unedited.
What is included
- Scorecard: accuracy, completeness, policy, tone and efficiency, each with anchored descriptions for every score point so reviewers are not inventing their own scale.
- Sampling plan: a risk-weighted selection method that prioritises refunds, policy exceptions, low sentiment, reopened threads and unedited AI sends.
- Calibration agenda: a 60-minute session format with facilitation notes and the rules that keep it useful, including scoring before discussion.
- AI supplement: the additional criteria for grading model-drafted replies and the quality of AI-to-human hand-offs.
- Reporting sheet: a simple monthly view including inter-rater agreement, which is the number that tells you whether your scores mean anything.
How to adopt it
Start with the scorecard and a small weekly sample; resist the temptation to grade everything. Run your first calibration session in week two, before habits form. Expect low inter-rater agreement initially, that is the normal starting point, and it improves quickly once disagreements are written up as rubric amendments.
Adapting it
The rubric is intentionally generic on tone, because tone standards are specific to your brand. Everything else transfers as-is across support organisations. If you operate in a regulated environment, add a compliance dimension rather than folding it into policy, so it can be reported separately.