Skip to main content
Call QA strategy

Call QA coverage vs accuracy: the new role of sampling

Full coverage helps find problem calls, but it does not prove a result is correct. This guide explains the new role of call center QA sampling: it shows the three numbers that separate coverage, fitness for analysis and accuracy with a worked example, lists rules for calls unfit for analysis, and sets out how to choose a sample that checks the system itself.

September 29, 20267 min read

The short answer

Analysing every call increases the chance of finding a problem call, but it does not make the evaluation any more correct. These are two separate questions. Coverage answers "how many calls were looked at"; accuracy answers "is the result given for a call right". There is a third number that often goes unmeasured altogether: how many of the analysed calls were fit for analysis at all.

Call center QA sampling — people listening to a small share of calls — does not disappear with full coverage. Its role changes: it used to exist to evaluate agents, and now it exists to check the system itself.

The old problem with sampling

In classic quality control a few calls per agent are listened to each month. The sample is small and often not random: calls with complaints, long calls or calls a team lead happens to remember get picked. The view of an agent then forms around their worst or most random days. A fuller treatment of that problem is in our guide to AI call quality audits.

AI analysis partly solves it: a system can look at hundreds of calls with the same criteria. But a new trap appears — "we looked at all of them" is heard as "we evaluated all of them correctly".

The illusion of full coverage

Not every analysed call is fit for evaluation. A call that went to voicemail, a recording where one side cannot be heard, a very noisy conversation, a ten-second call where the customer said "I'll call back" and hung up — all of these get analysed and all of them get some result. If those results flow into the overall statistics, both the average score and the comparison between agents are distorted.

The second illusion concerns "not applicable". If most criteria did not apply to a call, the score is based on very few of them. A call that met two out of two criteria gets a high score, but it is not as reliable as a call that met nine out of ten.

Three numbers, three questions

Show these three measures separately in a quality report. The figures below are an illustrative example, not a real company:

95%Coverage: 1,900 of 2,000 calls in the month were analysed
90%Fitness: 1,710 of the 1,900 analysed calls were fit for evaluation
85.5%Reliable coverage: 1,710 ÷ 2,000 — the share of calls whose result can be relied on

Accuracy is a third, separate number, and it comes only from a verification sample. Say a person has evaluated 120 calls independently: on "was a next step agreed" the AI reached the same result as the person on 108 calls (90%), and on "was the need clarified" on 96 calls (80%). That shows the second criterion needs to be written more precisely — and it still says nothing about the agents.

How to set aside calls unfit for analysis

Write the unfitness rules in advance, so a call is set aside on an objective sign rather than because of its evaluation result:

  • The call is shorter than the minimum the company has set and contains no real conversation
  • One side cannot be heard, or most of the speakers remained "unknown"
  • A significant part of the transcript is made of poorly recognised lines
  • The call reached voicemail, an answering machine or a wrong number
  • The conversation is in a language or call type the form was not designed for

An unfit call does not count against the agent and does not enter the denominator of the overall score. But it is counted separately: if the share of unfit calls on one shift suddenly rises, the problem may lie in telephony or recording settings.

How the verification sample should be chosen

With full coverage, the purpose of sampling is to check the system, so the sample has to cover all of its results, not only the bad ones:

  1. StratifyCreate separate groups by call type, shift and language, and pick calls from each.
  2. Cover the score bandsTake calls from low, middle and high scores — mistakes where the AI said "met" only show up among high-scoring calls.
  3. Evaluate independentlyThe person checking should not see the AI result first, or they will tend to agree with it.
  4. Compare by criterionCalculate agreement for each criterion separately, not across the whole form.
  5. Measure agreement between people tooHave two specialists evaluate the same calls. If people do not agree with each other, expecting agreement from AI is pointless.

There is no universal figure for sample size. Google's Quality AI guidance, for instance, asks for at least 100 example conversations per question to train its own model — that is a requirement of that product, not a rule for every system. The practical test is this: for every criterion there should be enough "yes" and "no" examples to see where the disagreements are.

Reliable coverage differs by agent

The three numbers differ by agent as well as by company, and that changes comparisons directly. An illustrative example: 285 of the 300 monthly calls of an agent on the day shift are fit for evaluation, while only 210 of a night-shift colleague's 300 are — the night line gets more noisy mobile calls and dropped connections. If both average 78, the first figure rests on 285 calls and the second on 210.

Two practical conclusions follow. First, an agent report should show next to the score how many fit calls it is based on. Second, the share of unfit calls should not be read as a measure of the agent: it is usually a feature of the shift, the line or the customer flow. For a fair comparison between agents, call type, shift and sample size always travel with the score.

Limits

High agreement does not prove the system is always right — only on the sample and in the period that were checked. When call types, scripts or products change, agreement has to be measured again. A person's evaluation is not automatically the "right answer" either, which is why agreement between people is tracked separately. Finally, the decision to set aside unfit calls should also be documented, so it is never used to hide inconvenient examples.

What Vexvon Audio Analyzer shows

Vexvon Audio Analyzer shows, per call, the signals these three numbers need:

  • Each analysis batch shows how many files completed and how many failed; a failed recording can be queued again without re-uploading
  • Poorly recognised lines are marked in the transcript, and a line whose speaker could not be determined stays "unknown"
  • A step of the standard that does not apply to the call gets "not applicable" and does not count against the agent
  • Every status comes with references to transcript lines, so the person checking compares the result with the evidence

The fitness rule, the verification sample and the agreement table are your quality process. The split of work between AI and people is covered in detail in automated call QA vs manual review.

First step

Add two lines to your next quality report: the share of calls fit for analysis, and AI–human agreement on at least two criteria. In the first month these numbers may look worrying — that is normal, because they were never measured before. More strategy pieces are in the call QA strategy section; to test on your own calls, get in touch.

Further reading on this topic: call recording quality speech analytics, speech analytics accuracy testing.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.