Audio vs transcript analytics: tone, empathy and text
Asking a transcript-based system "did the agent speak in a friendly voice?" is asking a question it cannot answer. This article separates what a call's text establishes reliably, what needs suitable audio and tooling, and what neither establishes reliably; it gives an illustrative boundary table and shows how to turn subjective criteria such as empathy, attentiveness and professionalism into behaviour visible in the text, with one sentence read two ways and five questions for a tool that promises audio analysis.
Short answer
Transcript analysis evaluates what was said in a call: which question the agent asked, which price they gave, what they promised, how they closed. Audio analysis tries to measure how it was said: tone of voice, pace, loudness, pauses. The first covers most of an agent's behaviour and is easy to check. The second needs suitable audio, specialised tooling and separate validation. And there is a third group: things neither text nor audio establishes reliably — for example, whether the agent was "really" empathetic.
Knowing this boundary matters when writing an evaluation form. Asking a transcript-based system "did the agent speak in a friendly voice?" is asking a question it cannot answer.
What text establishes reliably
- Stated facts: price, date, conditions, product name — provided the transcript is correct
- Whether questions were asked: "the agent asked about the budget"
- Whether mandatory statements were made: disclosure, confirmation, terms
- Promises and commitments: "we'll call you tomorrow", "your money will be refunded"
- The call's structure: greeting, clarifying the need, offer, close
- Prohibited phrases and wrong information given to the customer
Everything on this list is visible in words and can be evidenced with a transcript line. That is where most of agent evaluation lives.
What needs the audio
The following live in the sound itself, not in the words. Measuring them requires an audio analysis tool, recordings of sufficient quality and validation on your own calls:
- Tone and intonation — "I understand" can be said sincerely or sarcastically
- Speaking pace — a customer may not follow an agent who talks too fast
- Loudness and its changes — a raised voice
- The exact length of silences and holds
- How long overlapping speech lasted and who cut across whom
Google's Quality AI guidance states plainly that transcript-based evaluation does not perform acoustic analysis. AWS's evaluation documentation likewise advises avoiding questions about qualities only audio shows, such as loudness, tone and pace.
What neither establishes reliably
- The agent's inner feelings and intent — whether they "felt" empathy
- The customer's real mood — words and voice do not show their inner state precisely
- The agent's personality and character — no such conclusions should be drawn from voice or speaking style
- Protected characteristics — age, gender, ethnicity, health; these must not be inferred from voice or used in evaluation
- What happened after the call — a CRM entry, a refund, a promised callback
Illustrative boundary table
The table below is an illustrative example: typical criteria and the source they can be checked from.
- "The agent introduced themselves"Text: yes. Visible in words.
- "The agent gave the exact price"Text: yes, but numbers are sensitive to recognition errors — if it is a critical criterion, a person checks.
- "The agent spoke in a calm, friendly voice"Text: no. Audio: only with suitable tooling and validation. Better: turn it into behaviour.
- "The agent listened to the customer"Neither directly. Turned into behaviour: "restated what the customer said in their own words", "did not ask the same question twice".
- "The agent was empathetic"Neither reliably. Turned into behaviour: "acknowledged the complaint", "explained what was possible".
- "The customer was satisfied"Not from the call — from a customer survey or their later behaviour.
Turning a subjective criterion into behaviour
Subjective lines in a form do not need to be deleted — they need to be turned into behaviour visible in the text. The method is simple: ask "if this quality were present, what would the agent say?"
- EmpathyRestated the customer's problem in their own words; openly acknowledged the complaint.
- AttentivenessDid not ask again for information the customer had already given; did not forget the second question.
- ProfessionalismGave precise information; on a question they did not know, did not guess but said they would check.
- CourtesyDid not cut across the customer; asked at the end whether they had any other questions.
When audio analysis is needed
Sometimes the sound itself really does matter — for example, the exact length of a hold, or the agent talking over the customer. If you need such a criterion:
- Check in the documentation and on your own calls that the tool measures that metric
- Confirm that recording quality and channel layout are sufficient
- Compare the result with human evaluation
- Use the metric as a selection criterion for review, not an automatic score
Illustrative example: one sentence, two readings
The example below is illustrative. Two calls contain the same transcript line: the agent says "I understand, that really is frustrating."
- First callThe agent says it calmly, then asks for the order number and explains the new date. The customer says "fine, I'll wait" and ends the call.
- Second callThe agent says the same sentence sarcastically, with a sigh, and immediately adds "but those are the rules". The customer says "you're mocking me".
An evaluation reading the transcript will mark "the complaint was acknowledged" as met on both calls. The audio shows the difference — but the text also gives clues: on the second call there is no solution after the acknowledgement, and the customer's next sentence shows the displeasure growing. So writing the criterion as "acknowledged it and then explained what was possible" catches much more, even without measuring tone.
Five questions for a tool that promises audio analysis
- What exactly does this metric measure, and how is it computed?
- In which languages and at what recording quality has it been tested?
- Has the result been compared with human evaluation, and how?
- Does the result differ between stereo and mono recordings?
- How should the metric be used in decisions about agents — and how should it not be used?
Typical mistakes
- Expecting a transcript-based system to judge tone and intonation
- Accepting figures like an "empathy score" without knowing what they measure
- Drawing conclusions about an agent's character from their voice and speaking style
- Reading a sentiment label as a measure of tone of voice
Limits
- The transcript has errors of its own; facts taken from text depend on recognition quality
- Audio metrics are sensitive to recording quality, channel layout and language
- One tool's audio metric cannot be compared directly with another tool's
- Protected characteristics must not be inferred from voice or used in evaluation
What Vexvon Audio Analyzer offers
- The recording is transcribed, and each line is attributed to the agent, the customer or "unknown"
- The steps of your call standard are judged from the transcript as met, partial, missed or not applicable, with evidence lines
- When the standard changes, evaluation is re-run on the existing transcript — in other words, the evaluation rests on the text
- No separate metrics for tone of voice, intonation or speaking pace are offered
More: Vexvon Audio Analyzer. Reading silences and interruptions is covered in call silence and interruptions, and the limits of sentiment in customer sentiment vs agent performance.
First step
Open your evaluation form and write next to each line: text, audio or neither? Turn the "neither" lines into behaviour using the method above. This article belongs to the audio analysis implementation and reliability section. To try it on your own recordings, get in touch.
Further reading on this topic: speaker diarization call center.