QA scorecard versioning: comparing scores after criteria change
When a criterion gets stricter the average falls, yet agents are talking as before. This guide explains QA scorecard versioning: which changes need a new version, a version record, a golden set of fixed test calls, the steps of the transition period, and why a model change is a version too.
The short answer
When evaluation criteria change, new and old scores are not directly comparable, because they were measured with different rulers. The right approach is QA scorecard versioning: every change gets a new version number, every call keeps the version it was evaluated against, during the transition the same test calls are evaluated with both versions, and results that cannot be compared are clearly marked in the report. That separates "quality fell" from "the ruler changed".
What needs a new version
- The criterion's text, or the definitions of "yes", "partial" and "no"
- The "not applicable" condition
- Adding or removing a criterion
- Weights or the scoring rule
- The call types the form applies to
- The AI model that runs the evaluation, or its instructions — even if the form stays the same
A typo fix does not need a version. Adding "a time must also be stated" to the "next step" criterion is already a different ruler.
Why old and new scores are not comparable
An illustrative example. A team tightens the "next step" criterion: "we'll call you back" used to count as "yes", and now a specific time or responsible person is required. The following month the average drops by five points. Management concludes that "quality has got worse" and plans training. Yet agents are talking exactly as before — only the ruler has changed.
AWS's evaluation-forms documentation flags the same problem for scoring modes: it recommends creating a new form rather than switching the scoring mode on an existing one, because evaluations completed under the old mode cannot be directly compared with the new one.
A version record: template
- Version number and effective date
- What changed — the criterion, with the old and new text side by side
- Why it changed — a calibration decision, an appeal, a new product, a legal requirement
- Expected effect — on which criterion scores will rise or fall
- The actual effect measured on the test calls
- The decision about the old period — re-evaluated or shown separately
- Who approved it and when it was announced to agents
A "golden set": the same test calls
Keep a fixed set of 20–30 calls: of different types, with results agreed by people. Every new version is checked on this set first. It shows how much the new version changes the ruler, independently of behaviour.
- Golden set, v3The 25 calls average 74.
- Golden set, v4The same 25 calls average 69 — the ruler has become 5 points stricter.
- Real month, v3Last month's average: 76.
- Real month, v4This month's average: 71 — a difference of −5.
- Change in behaviour(71 − 76) − (69 − 74) = −5 − (−5) = 0. Agents' behaviour has not changed; the drop comes entirely from the new ruler.
The transition step by step
- Parallel evaluationBefore the new version takes effect, the golden set and some calls from the last two weeks are evaluated with both versions.
- Measure the effectThe difference per criterion is calculated and written into the version record.
- Announce itAgents are shown the date, the changed criterion and an example of it.
- Draw a line in the reportThe version change is marked with a line or note on the chart; old and new averages are not joined into one line.
- Decide about the old periodIf a trend is needed, some old calls are re-evaluated with the new version; if not, the periods are shown separately.
A model change is a version too
Even if the form stays the same, the AI model that runs the evaluation, or its instructions, can change — and so will the results. AWS handles this separately: its forms store the generative-AI version with the form version, and you can pin a specific version or try a new one ahead of time. Whatever tool you use, the principle is the same: record a model change as a version and check it on the golden set.
Announcing it to agents: a sample text
Agents should hear about a version change in advance, not from a report. An illustrative announcement: "From 1 November the 'Next step' criterion changes. 'We'll call you back' is no longer enough — you need to give a specific day or time, or the name of the person responsible. Example: 'Our manager Rauf will call you tomorrow at 11.' After this change your scores on this criterion may look lower for a while; they will not be compared directly with October's results."
The last sentence matters: it tells agents in advance that a lower score will not be read as their work getting worse.
Versioning in a small team
For a team of ten to fifteen agents the whole process can seem heavy. The minimal version: a version number and date at the top of the criteria document, a one-page change list and a ten-call golden set. The smaller the golden set, the more carefully you read its result — but even ten calls give a rough idea of how much the ruler moved and head off a "quality has dropped" panic.
Limits
A golden set is small and may not cover every call type — the ruler's effect in a real month can differ a little. The estimate of behaviour change is approximate, especially if the call flow changed at the same time. Changing versions too often makes the trend unreadable: batch criterion changes and make them once or twice a quarter, except for urgent legal requirements.
Versions in Vexvon Audio Analyzer
Vexvon Audio Analyzer is built for this approach. Every time the standard's text changes, the version number goes up. Each call keeps which version of which standard it was evaluated against, along with a copy of the text at that moment — editing or deleting the standard later does not change what an old score meant. An existing transcript can be re-evaluated with a new version without going back to the audio, so re-scoring the golden set for every version is not expensive.
The reason for a version is usually a calibration decision; for comparison rules, see fair agent performance comparison.
First step
Build a golden set from 20–30 calls with human-agreed results and record your current form's result on it. Before the next criterion change, evaluate the same set with the new text. More in the agent scoring section; get in touch.
Further reading on this topic: call center scorecard template.