Skip to main content
Audio analysis

Call recording quality: how poor audio distorts scores

If an agent on a badly recorded line scores lower than everyone else every month because of a microphone or the network, the evaluation system is measuring the equipment, not the agent. This article covers five typical recording problems in contact centers, how an error travels from the transcript to the score, the difference between "no", "not applicable" and "unable to assess", an illustrative recording-quality rule, showing score coverage, a recording-problem log and fixing the cause.

September 29, 20267 min read

Short answer

A poor recording distorts an agent's score in two ways: the system does not hear a word and the agent is marked as not having said it, or it mishears a word and attributes to the agent something they never said. Either way the fault lies with the recording, not the agent. So a poor or incomplete recording must not automatically produce a negative score — on such a call the criterion should be marked "unable to assess", not "no".

This is a matter of fairness, not a technical detail. If an agent working on a badly recorded line scores lower than everyone else every month because of a microphone or the network, the evaluation system is measuring the equipment, not the agent.

Five typical recording problems

  1. Background noiseAn open office, a street, a car, a TV. Short words — "yes", "no", numbers — are lost most, and those are exactly the confirmations and prices.
  2. Dropouts and packet lossOn a weak network, part of a sentence never reaches the recording. In the transcript the sentence is cut off or two sentences merge.
  3. ClippingWhen audio is recorded too loud, the peaks of the waveform are cut and the sound crackles. Recognition errors rise, especially with a customer who speaks loudly.
  4. Incomplete recordingThe recording starts mid-call or stops before the end. If the greeting or the close is missing, those criteria cannot be assessed.
  5. A low-quality codecA heavily compressed format can keep speech understandable to a person but reduce automatic recognition accuracy.

How an error turns into a score

An error from recording quality spreads like a chain. The example below is illustrative:

  1. RecordingThe agent says: "Your order will be delivered on Thursday, between 2 and 6 pm." At that moment the line drops for two seconds.
  2. Transcript"Your order will be delivered on Thurs ... 6 pm." The day and time window are cut off.
  3. EvaluationThe criterion "the agent stated the delivery time precisely" comes out as no or partial.
  4. ResultThe agent's score drops, and the feedback conversation shows them a mistake they never made.

A checkpoint can be placed on every link of that chain: in the recording itself, in the transcript's reliability and in the evaluation status.

Three statuses: no, not applicable, unable to assess

  1. NoThe criterion applied, the recording is clear enough, and the behaviour is absent. Only this status lowers the score.
  2. Not applicableThe criterion does not belong to this call: for example, the customer did not object, so objection handling is not checked. Left out of the score.
  3. Unable to assessThe criterion could have applied, but that part of the recording is lost or unreliable. Left out of the score and counted separately in the report.

Without the third status, the system is stuck between two bad choices: say "no" and penalise the agent, or say "not applicable" and hide the problem. Both mislead the report.

Illustrative recording-quality rule

The rule below is an illustrative example. Set the thresholds from your own recordings and pilot results.

  1. Full evaluationThe recording covers the call from start to end, low-confidence lines are few and not at moments that matter for the criteria.
  2. Limited evaluationPart of the recording is poor. Criteria that fall in that part get "unable to assess"; the rest are evaluated normally. The score is shown with its coverage.
  3. No evaluationThe recording is very short, one side's audio is missing, most of the conversation is unintelligible. The call gets no score, is logged as a recording problem and passed to the technical team.
  4. Human reviewIf a critical criterion — such as a mandatory disclosure — falls in a poor section, a person listens to the audio and makes the call.

Showing coverage

If a score is computed only from the criteria that could be assessed, the number of assessed criteria must be shown next to it. The example below is illustrative:

  • Call A: 10 of 10 criteria assessed, score 80
  • Call B: 5 of 10 criteria assessed, score 80 — the other 5 could not be assessed because of recording quality
  • The same "80" carries different reliability; adding B to the agent's average with the same weight as A is wrong

The same rule applies at agent level: if the share of "unable to assess" calls is high, the agent's average is unreliable and the technical state of their line should be checked.

Fixing the cause

  • Check headsets and microphones — if all of one agent's calls are poor, the cause is often the equipment
  • Adjust recording levels — audio that is too quiet is not recognised, too loud is clipped
  • If possible, record the agent and the customer on separate channels
  • Avoid over-compressed formats; balance storage savings against recognition quality
  • Check that the recording covers the whole call — including after a transfer

A recording-problem log

Recording-quality problems are rarely random — they repeat on the same line, the same equipment, at the same hours. A simple log is enough to see them. One row is written for every "unable to assess" or "no evaluation" case:

  1. Date and lineWhen the call happened and on which line or shift.
  2. Problem typeNoise, dropout, clipping, incomplete recording, one side's audio missing.
  3. How it was foundLow-confidence lines, a failed analysis, a person listening, or the agent disputing a result.
  4. ImpactWhich criteria could not be assessed and how that might have affected the agent's score.
  5. Action takenHeadset replaced, level adjusted, passed to the technical team.

After a month, the log shows where problems cluster and gives a concrete basis for technical investment: "six desks on the third floor have a higher dropout share than the rest".

Typical mistakes

  • Counting a criterion missed on a poor recording as "no"
  • Quietly dropping poor recordings from the sample without reporting it — problem lines become invisible
  • Comparing agents' averages without looking at recording quality
  • Reading the recogniser's confidence value as a calibrated probability

Limits

  • Per-line recognition confidence is the model's own estimate; it is not a calibrated probability and does not guarantee the line is correct
  • A transcript that looks clean can still be wrong — especially numbers, names and product names
  • The thresholds of a recording-quality rule must be tested on each company's own recordings
  • A call that cannot be assessed because of a poor recording must not count against the agent

What Vexvon Audio Analyzer offers

  • Each transcript line comes with the recogniser's confidence; low-confidence lines are marked openly in the panel — not hidden
  • A line whose speaker cannot be determined is kept as "unknown", not guessed
  • A file that fails analysis is shown with the error reason and can be re-queued without uploading it again
  • A call-standard step marked not applicable does not count against the agent; the "unable to assess" rule above is applied in your own process, with human review

More: Vexvon Audio Analyzer. How to test recording quality in a pilot is shown in the speech analytics pilot plan, and the recording fields in call recording metadata fields.

First step

Listen to last week's 10 lowest-scoring calls and ask one question of each: does the low score come from the agent's behaviour or from the recording? List the recording-related ones separately and check which line and which equipment they belong to. This article belongs to the audio analysis implementation and reliability section. To try it on your own recordings, get in touch.

Further reading on this topic: not applicable in call QA scoring, speaker diarization call center, speech analytics accuracy testing.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.