Skip to main content
Audio analysis

Speaker diarization in call centers: why it matters

The customer says "you'll give me this for free", the transcript puts that line under the agent, and the agent loses points for an unconfirmed promise they never made. In the transcript the error looks completely normal. This article covers what speaker diarization is, the difference between stereo and mono recordings, four consequences of misattribution, the most sensitive criteria, an illustrative transcript excerpt, a test method, a rule for "unknown" speakers and ways to reduce errors.

September 29, 20267 min read

Short answer

Speaker diarization is the process of separating the voices on one audio channel and attributing each sentence to a speaker. For agent scoring it is a decisive step: if the system attributes the customer's sentence to the agent, the agent is scored on words they never said. If the customer said "this is illegal" but the transcript puts it under the agent, the agent looks as if they made a serious mistake.

The most reliable way to reduce that risk is to record the agent and the customer on separate audio channels. Where that is not possible — on mono recordings — diarization quality needs checking on its own, and lines the system is unsure about should not be guessed.

Stereo and mono: two different situations

  1. Stereo — separate channelsThe agent's microphone goes to one channel and the customer's line to the other. Who spoke is known physically. Errors are rare; the main risk is a wrong channel mapping — for example, the left channel is the agent on one PBX and the customer on another.
  2. Mono — one channelBoth voices are recorded mixed on one channel. The system groups speakers by voice characteristics and then decides which group is the agent. Errors are possible at both steps.
  3. Three or more voicesA warm transfer, a conference, a supervisor joining. Separating three voices on a mono recording is harder. AWS's generative evaluation guidance also notes that overlapping speech and multi-party calls reduce accuracy.

What misattribution leads to

A wrongly identified speaker shows up in evaluation in different ways. The cases below are an illustrative example:

  1. A false violationThe customer says "you'll give me this for free", and the transcript shows it as the agent's line — the "unconfirmed promise" criterion counts as broken.
  2. A false successThe customer describes their need in detail and that part is attributed to the agent — "the need was clarified" counts as met, although the agent asked nothing.
  3. A lost mandatory statementA mandatory disclosure the agent made is attributed to the customer and counted as not said.
  4. Distorted interruptionsShort acknowledgements land on the wrong side and the interruption count becomes meaningless.

Which criteria are most sensitive

  • Criteria tied to a specific thing the agent said: a mandatory disclosure, a price, a promise, a condition
  • Criteria tied to asking: "the agent asked about the need" — who asked matters
  • Prohibited phrases: if the customer's words are attributed to the agent, a false violation appears
  • Interruptions and talk share — both depend on identifying the speaker correctly

General criteria — such as "the next step was clear at the end of the call" — are usually more robust, because one or two misattributed lines do not change the result.

Illustrative check: how to test diarization

The method below is an illustrative example. The goal is not an overall percentage but to see whether errors happen at the moments that matter for evaluation.

  1. SampleChoose 10–20 mono calls across different agents and customers; include one or two three-party calls.
  2. Key linesIn each call, mark the 5–10 lines that matter for evaluation: greeting, price, promise, mandatory statement, close.
  3. Listen and checkFor each marked line, listen to the audio and note whether the speaker was identified correctly.
  4. ResultThe share of misattributed key lines and the share left as "unknown". When do errors happen: short sentences, overlapping speech, similar voices?

Working with an "unknown" speaker

When the system is unsure, keeping a line as "unknown" is more honest than guessing it onto the agent or the customer. But such lines need their own rule in evaluation:

  • If the evidence for a critical criterion is only in an "unknown" line, a person listens to the audio and decides
  • A call with a high share of "unknown" lines is not treated as fit for full evaluation
  • If one line or one agent consistently has a high "unknown" share, the cause is in the recording setup
  • An "unknown" line is not used as evidence against the agent

Reducing diarization errors

  1. Record on separate channelsIf the PBX and recording system allow it, this is the most effective step. Check storage volume and consent rules alongside it.
  2. Check the channel mappingOn stereo recordings, confirm with a few calls which channel is the agent; lines and PBXs can differ.
  3. Audio qualityNoise and dropouts make diarization worse too. Headset and microphone quality matter here as well.
  4. Three-party callsFlag warm transfers and conference calls separately and apply more human review to them.

Illustrative transcript excerpt

The excerpt below is an illustrative example, not a real call. On a mono recording, diarization has put two lines on the wrong side. The note on each line shows the correct speaker after listening to the audio.

  1. Agent: "The plan is 25 a month, with a discount in the first month."Correct — the agent said it.
  2. Agent: "So the first month is completely free?"Wrong — this is the customer's question. The system mixed up the voices.
  3. Customer: "No, the first month is 50% off, so 12.50."Wrong — this is the agent's answer.
  4. Customer: "Got it, let's sign up then."Correct — the customer said it.

An evaluation that reads this excerpt from the transcript will see the agent making an unconfirmed promise that "the first month is free" and the customer stating the exact price. The result is completely reversed: the agent did the job right and still loses points. That is why the evidence lines for critical criteria must be checked by listening to the audio.

Four checks before talking to the agent

  • Was the speaker of the line showing the violation confirmed by listening to the audio?
  • Is the recording stereo? If mono, was there overlapping speech at that moment?
  • Is the evidence line not marked "unknown"?
  • Are other lines in the same call also mixed up — is the whole call reliable?

Typical mistakes

  • Trusting the transcript's speaker labels without checking
  • Comparing mono and stereo results as if they were equally reliable
  • Discussing a false violation with the agent in a feedback conversation without listening to the audio
  • Quietly assigning "unknown" lines to the agent or the customer

Limits

  • Diarization accuracy varies with language, audio quality, number of speakers and how similar the voices are; a result measured in one setting does not transfer to another
  • Stereo recording reduces errors, but if the channel mapping is wrong the whole call is reversed
  • During overlapping speech, words from both sides can be lost or mixed
  • A result caused by a wrongly identified speaker must not be used against the agent

What Vexvon Audio Analyzer offers

  • On a stereo recording the agent and customer are separated by channel; on mono, by diarization
  • When the system is unsure, the line is kept as "unknown" — not guessed onto the agent or the customer; each line's speaker is visible in the transcript
  • The diarizer's original label is stored — if the agent/customer mapping is wrong, it can be traced
  • Each step of the call standard comes with the transcript lines behind it — so you can check directly whom the evidence line was attributed to

More: Vexvon Audio Analyzer. Why channel layout matters as metadata is covered in call recording metadata fields, and recording quality in general in call recording quality and scoring.

First step

Check whether your recordings are stereo or mono. If mono, use the method above on 10 calls and listen to confirm the speaker of each key line. This article belongs to the audio analysis implementation and reliability section. To try it on your own recordings, get in touch.

Further reading on this topic: evidence-based call scoring, speech analytics accuracy testing.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.