Skip to main content
Audio analysis

Speech analytics accuracy testing: beyond one number

"Our system is 90% accurate" is not useful information unless it says what was measured, on which calls and in which language. This article splits AI call analysis accuracy into four separate measures, uses an illustrative calculation to show how 89% agreement can hide a quarter of real violations being missed, and covers building a reference set, why a confidence value is not a probability, cause analysis for disagreements, a one-page accuracy report and monitoring accuracy over time.

September 29, 20267 min read

Short answer

The accuracy of AI call analysis cannot be measured with one number. At least four things need measuring separately: transcript quality, per-criterion agreement between AI and a human reference, the type of error on critical criteria — false flags and missed violations — and the share of calls that cannot be analysed. "Our system is 90% accurate" is not useful information unless it says which of the four it measured, on which calls and in which language.

In other words, accuracy is not a property of the tool; it is the result the tool achieves on your calls, against your criteria. It can only be measured on your own reference set.

Four separate measures

  1. Transcript qualityHow correctly speech was turned into text. An overall word error rate helps, but for evaluation the key words matter more: price, date, product name, mandatory phrase.
  2. Per-criterion agreementFor each criterion, the share of calls where the AI answer matches the reference. Computed criterion by criterion, because one may work well and another poorly.
  3. Error typeOn a critical criterion the two errors have different consequences: a false flag accuses the agent unfairly, a missed violation hides a risk. They are counted separately.
  4. Unanalysable callsThe share of calls that cannot be scored reliably because of audio quality, language or speaker separation. These are removed from the agreement count and shown separately.

Illustrative example: what 89% agreement hides

The figures below are an illustrative example, not any tool's real results. Critical criterion: "the agent made the mandatory disclosure". Of 108 calls, 8 could not be analysed because of recording quality; the remaining 100 were compared with a human reference. According to the reference, the disclosure was made on 80 calls and not made on 20.

  1. The 80 calls with the disclosureAI answered "made" on 74 and "not made" on 6. Those 6 are false flags.
  2. The 20 calls without itAI answered "not made" on 15 and "made" on 5. Those 5 are missed violations.
89%Overall agreement: (74 + 15) / 100
25%Share of real violations missed: 5 / 20
29%Share of flags that were false: 6 of 21

89% looks good. But a quarter of the real violations were not found, and roughly one in three raised flags was wrong. On top of that, 8 of 108 calls were not scored at all — coverage is 92.6%. The report should show all four figures together, not just 89%.

Building the reference set

  • The calls reflect real traffic: different agents, call types, languages and recording quality
  • Critical criteria have enough "yes" and "no" cases — rare violations may need targeted selection
  • Two people score each call independently; differences are resolved by discussion
  • The reference set is kept stable, so that when a criterion or model changes you can compare on the same calls
  • The people building the reference do not see the AI result beforehand

A confidence value is not a probability

Many systems give a confidence value for a transcript line or a decision. It is a useful signal — low-confidence lines deserve a closer look. But "0.9" does not mean that 90% of such decisions are correct. To read it that way, the value has to be calibrated on your calls, that is, checked against the reference.

A practical approach: use confidence for sorting and selection, not as a probability. On the reference set, check whether low-confidence decisions really do contain more errors.

Finding the cause of disagreement

For every case where AI and the reference differ, the cause is recorded in one category. The list shows what needs fixing:

  1. Transcript errorA word was misrecognised or lost. Fix: recording quality, a terminology glossary.
  2. Speaker errorA line was attributed to the wrong side. Fix: separate-channel recording, a diarization check.
  3. Criterion wordingThe criterion can be read two ways. Fix: tighten it and add "yes" and "no" examples.
  4. Interpretation gapThe transcript and criterion are clear, but the AI reached a different conclusion. Fix: more examples, and keep that criterion under human review.
  5. Reference errorOn review it turns out the person made the mistake. The reference is corrected — people are not flawless either.

Monitoring accuracy over time

Accuracy measured once changes over time: a new product, a new script, new agents, a new customer mix. AWS's generative evaluation guidance also notes that AI evaluations are not 100% accurate, that samples should be reviewed by people regularly, and that a manual evaluation process should stay in place to catch drift.

  • Each month a small sample is re-scored by people
  • When a criterion or call standard changes, the reference set is re-checked with the new version
  • If agreement drops on a critical criterion, that criterion temporarily returns to human review
  • Agent disputes are a signal too — which criteria do the upheld disputes cluster on?

A one-page accuracy report

Accuracy information shown to management and agents should be short but complete. The structure below is an illustrative example:

  1. ContextWhich call type, which languages, which period, how many calls in the reference set, which version of the call standard.
  2. CoverageHow many calls could be analysed, how many could not, and the main reasons.
  3. Criterion tableAgreement for each criterion; for critical criteria, also the share of missed violations and false flags.
  4. Usage ruleWhich criteria are used as automatic results and which go through human review.
  5. Next checkWhen, and after which change, the measurement will be repeated.

Sharing this report with agents builds trust: they know where the system is strong and where it is weak, and can tie their disputes to a specific criterion.

Typical mistakes

  • Assuming a vendor's overall accuracy figure also holds for your calls
  • Computing only overall agreement without separating error types
  • Quietly dropping unanalysable calls and not reporting coverage
  • Building the reference while looking at the AI result
  • Quoting an old accuracy figure after a criterion has changed

Limits

  • Measured accuracy applies only to that reference set — call type, language, recording quality
  • A reliable figure for rare violations needs many calls; a percentage from a small sample is read with caution
  • A model's confidence value should not be presented as a probability unless it is calibrated
  • No accuracy level makes an AI result the sole basis for a decision about an agent

What Vexvon Audio Analyzer offers

  • Each transcript line comes with recognition confidence; low-confidence lines are marked
  • Each step of the call standard has a status, a comment and evidence lines — the information you need to compare with the reference and find the cause of disagreement
  • After fixing criteria, the same reference calls can be re-analysed from the existing transcript — for a direct comparison of results
  • No overall accuracy percentage is offered: accuracy should be measured on your reference set

More: Vexvon Audio Analyzer. How to build a reference set in a pilot is covered in the speech analytics pilot plan, the effect of recording quality in call recording quality and scoring, and per-language testing in multilingual call QA.

First step

Pick one critical criterion and build a reference set of 50–100 calls. Compare the AI result with it and write down four figures: agreement, missed violations, false flags and unanalysable calls. This article belongs to the audio analysis implementation and reliability section. To try it on your own recordings, get in touch.

Further reading on this topic: AI call scoring calibration.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.