AI call scoring calibration: when the AI and QA disagree
When the AI and a quality specialist disagree, neither is automatically right. This guide builds AI call scoring calibration step by step: the flaws of human scores, how a session runs, an illustrative difference table, five causes of disagreement and their fixes, the agreement rate that misleads on rare events, and how often to calibrate.
The short answer
If the AI and a quality specialist score the same call differently, neither should automatically be taken as right. The right step is AI call scoring calibration — evaluating the same calls independently, comparing the differences criterion by criterion, finding the cause of each, and fixing the criterion or the process. The cause is often in neither the AI nor the person but in the criterion: the two sides read it differently.
The outcome of calibration is not "who won", but a clearer criterion, a documented decision and fewer differences at the next check.
A human score is not flawless either
A quality specialist's evaluation is usually taken as the "right answer", but people have systematic errors too. Fatigue softens criteria at the end of the day; a good or bad impression at the start of a call spills over into the other criteria; an evaluator who knows the agent may treat them differently. Two experienced specialists can score the same call differently.
Google's Quality AI guidance asks for several human evaluators to answer consistently on the examples used to train its model, because inconsistency introduces noise. The principle applies to any system: if people do not agree with each other, expecting AI to agree with them is pointless.
A calibration session step by step
- PreparationPick 10–15 calls: different types, different scores, at least a few disputed cases. Send everyone the criteria's definition cards.
- Independent evaluationAt least two specialists evaluate the calls without seeing the AI result.
- Difference tableFor each call and criterion, three results are written side by side: the AI, the first and the second specialist.
- Discussing causesOnly the lines that differ are discussed. For each difference a cause category is chosen (below).
- Decision and recordThe correct reading of the criterion is written down; the definition card is changed where needed and recorded as a new version.
- Re-testThe same calls are evaluated again with the new criterion text. If the differences shrink, the fix worked.
An illustrative difference table
A five-line example — not real calls:
- Call 2, "next step": AI — no, QA1 — yes, QA2 — yes. Cause: the agent said "I'll call you tomorrow at 11", but the line was attributed to the customer — a speaker error
- Call 4, "price conditions": AI — yes, QA1 — partial, QA2 — no. Cause: the criterion never says what counts as a "condition" — criterion wording
- Call 7, "identity confirmation": AI — not applicable, QA1 — no, QA2 — no. Cause: the AI misjudged the call type — an AI error
- Call 9, "clarifying the need": AI — yes, QA1 — yes, QA2 — no. Cause: QA2 did not count one question as enough — disagreement between people
- Call 12, "price is correct": AI — yes, QA1 — no, QA2 — no. Cause: the price changed that day, which the AI cannot know — outside information
Five causes of disagreement and their fixes
- Criterion wordingThe most common cause. Fix: add the "yes", "partial" and "not applicable" conditions and non-standard examples to the definition card.
- Transcript or speaker errorFix: check the data, not the criterion — recording quality, whether stereo recording is possible; put such calls in the review queue.
- AI errorThe criterion is clear and the transcript is right, but the result is wrong. Fix: split the criterion into simpler questions; if it persists, keep human review on that criterion.
- Disagreement between peopleFix: joint training for the specialists and a sharper definition card — no need to change the AI.
- Outside informationThe result depends on information that is not in the transcript. Fix: mark the criterion as "requires an outside check" and have only people evaluate it.
How to measure agreement: a misleading percentage
The simplest measure is agreement per criterion: on how many calls the AI and a person reached the same result. But for rare events that percentage misleads. An illustrative example: in 38 of 40 calls the agent greeted the customer. The AI wrote "yes" on all 40. Agreement looks like 38 ÷ 40 = 95% — yet the AI caught neither of the two calls where the greeting was missed.
When to calibrate
- Before launch — before the form is announced to agents
- Once a month for the first three months
- After every criterion change — testing the new version on old calls
- When a new call type, product or language is added
- When agents' appeals suddenly rise on one criterion
The calibration record: what to write down
A session's value lives in its record. Keep a short record in the same format for every decision:
- Date, participants and the list of calls reviewed
- The criterion and its text before the session
- The cause category of the disagreement and an example call
- The agreed reading — one sentence, for future evaluators
- The criterion's new text and version number
- Impact: which past results need another look, and whether agents will be told
Such a record is the best training material for a new specialist, and when an agent appeals it shows why the criterion is read the way it is.
Limits
Calibration is done on a small sample, and its result holds only for those call types. High agreement on one criterion says nothing about another. If a session's decisions are not written down, the same argument comes back three months later. And if calibration changes an agent's result, the agent should be told — especially if the result has already been discussed with them.
Calibration in Vexvon Audio Analyzer
Vexvon Audio Analyzer makes several steps of calibration simpler. Every status comes with evidence lines, so discussing a difference moves from opinion to facts. The raw speaker label is kept in the transcript, so a wrong separation can be traced. Most importantly: after the criterion text is fixed, an existing transcript is re-evaluated without reprocessing the audio, and each call records the version of the standard it was judged against. The "re-test" step takes minutes rather than days.
The session itself — who evaluates, the difference table, the decision record — is your process. For how evidence is structured, see evidence-based call scoring; for the verification sample, call center QA sampling.
First step
Run your first ten-call session next week: two specialists, the AI result and a difference table. Assign one of the five causes to every difference and start with the one that recurs most. More in the agent scoring section; to prepare the session together, get in touch.
Further reading on this topic: QA scorecard versioning, speech analytics accuracy testing.