Why did the AI give this score? Evidence-based call scoring
"This call scored 64" tells an agent nothing. This guide explains the four parts of evidence-based call scoring, an illustrative three-criterion example, how to evidence a behaviour that did not happen, how to check the evidence itself, the difference between a timestamp and a transcript line, and how to show evidence to the agent.
The short answer
In evidence-based call scoring each criterion's result has four parts: the criterion itself, the status (yes, partial, no, not applicable), a one- or two-sentence rationale, and the evidence that shows the result — a quote from the transcript and where it sits. The score is the sum of these records. When an agent or a manager asks "why this score?", the answer should be a specific line, not an opinion.
An AI score without evidence does damage in two ways: the agent does not believe it, and the manager never finds out when it is wrong. Evidence solves both at once.
Why a score without evidence does not work
"This call scored 64" tells the agent nothing. They do not know what they did wrong, so they argue with the result or ignore it. The team lead is in a difficult spot too: to talk to the agent they have to listen to the call from the start — redoing exactly the work AI analysis was supposed to save.
The more serious problem is the invisible mistake. If the AI mixed up the speakers and attributed the customer's words to the agent, a score without evidence will never show it. With evidence, the reviewer sees at a glance that the quote belongs to the customer.
The four parts of an evidence-based result
- CriterionA full-sentence question: "Did the agent agree a specific next step?"
- StatusYes, partial, no or not applicable — by the rule set on the form.
- RationaleWhy this status: "The agent promised a callback, but gave no time or responsible person." One or two sentences, in the reviewer's language.
- EvidenceA quote and its place in the transcript: "Line 14, agent: We'll call you back." The reviewer should be able to jump there in one click.
An illustrative example: three criteria
Not a real call — an example built for explanation. A customer wants to raise their credit card limit.
- Identity confirmation — Yes. Rationale: before giving any information the agent asked for the date of birth and the card's last digits. Evidence: lines 3–5
- Explaining the conditions — Partial. Rationale: the new limit was stated, but not whether the interest rate changes. Evidence: line 9, agent: "Your limit will go up to 3,000 manat"
- Next step — No. Rationale: nothing was said about when the decision will be made or how the customer will be told. Evidence: lines 12–14, the end of the call; agent: "We'll look into it and let you know"
In this format a team lead can prepare for the conversation without listening to the call: two behaviours, three lines. And if the agent disagrees with a result, they appeal by pointing to a specific line.
Evidence for a "no"
The hardest case is evidence of a behaviour that did not happen. How do you show "the agent did not state the cancellation condition"? Writing "not found" is not enough. A good record shows two things: where the behaviour was expected (the lines right after the agreement, for example) and what was actually said there. The reviewer then reads the expected window, not the whole call.
How to check the evidence itself
- Is the quote really in the transcript — or has the model "rewritten" it in its own words
- Is the speaker identified correctly — does the quote belong to the agent, not the customer
- Context — was the sentence taken back or corrected on the next line
- Recognition quality — if the line is marked as poorly recognised, listen to the audio itself
- Does the quote really support the status — the word "hello" does not prove "stated the company name"
A timestamp or a transcript line
Some systems show evidence as a time in the audio: in AWS's generative-AI evaluation you can choose the time associated with a transcript reference and jump to that point in the conversation. That is useful when the recording and the transcript times line up exactly. But when the time is the model's estimate and it drifts, a click lands somewhere else in the call and the reviewer listens to the wrong part. An unreliable time pointer is worse than none.
A reference to a transcript line avoids that risk: the evidence sits in the very text being checked. When the audio does need to be heard, the text around the line shows which part of the recording to look for.
How to show the evidence to the agent
Evidence should be shared before the conversation with the agent, not during it — so the agent has time to read the line and think. Keep the conversation to one or two criteria, not the whole form. And give the agent a way to dispute every result: if the line shown as evidence is wrong, there must be an open way to say so.
What evidence gives a manager
On one call, evidence makes the conversation specific; across hundreds of calls it lets you see patterns. Read the evidence lines of every call that got "no" on the same criterion side by side, and the same phrase often turns up — "we'll look into it and let you know", for example. That is a team habit, not one agent's, and the fix is in training or the script.
Evidence also exposes the system's own mistakes. If the evidence lines on one criterion often belong to the customer, the problem is speaker separation, not the agents. If the quotes do not support the criterion, its wording needs to be sharpened. Without evidence neither signal would be visible.
Limits
A rationale written by a model can sound convincing and still be wrong — which is why the evidence line matters more than the rationale. A recogniser's per-line confidence is not a calibrated probability. Evidence shows behaviour, not intent: why the agent spoke that way comes out in the conversation. And no evidence-based score should be the only basis for a decision about an agent.
Evidence in Vexvon Audio Analyzer
Vexvon Audio Analyzer is built around this model:
- Every step of the standard gets a status (met, partial, missed, not applicable) and a comment
- Evidence is a reference to transcript lines; the panel highlights those lines and scrolls to them
- Every transcript line is shown as agent, customer or "unknown"; a poorly recognised line is marked
- Step names, comments and recommendations can be translated into chosen languages
- The audio can be played in the panel, but evidence is not given in seconds — time estimates were unreliable, so they were removed on purpose
For what the statuses mean, see no, not applicable and unable to assess; for how criteria are written, the scorecard template.
First step
Take five evaluations from last week and check every "no": is there a rationale, is there a quote, does the quote really belong to the agent? The share of results without evidence shows how explainable your evaluations are. More in the agent scoring section; get in touch.
Further reading on this topic: AI call scoring calibration, call center QA appeals.