AI call quality assurance: from sampling to all calls
The classic form of call quality control is listening to a few calls a week. In a team making a thousand calls that is coverage under one per cent, and the sample is not random — leads tend to listen to the calls that produced complaints. This article turns AI call quality assurance into a system: how quality becomes measurable criteria, how the scorecard is written, what automation can and cannot do, the legal side of recording and consent, how findings get back to the team, and how the system gets misused.
Why listening to a sample is not enough
The classic form of call quality control is this: a team lead listens to a few calls a week and takes notes. It works, on one condition — that the number of calls is small. In a team making a thousand calls a week, the share a lead can listen to is under one per cent.
A one per cent sample can find a problem; it cannot measure a system. Worse, the sample is not random. Leads tend to listen to the call that produced a complaint or the one that ran long, so the data always leans towards trouble and how the ordinary calls went stays invisible.
This article is about turning quality review into a system: how quality becomes measurable criteria, how a scorecard is written, what automation can and cannot do, the legal side, and how findings get back to the team.
What quality means — measurable criteria
«A good call» cannot be measured. What can be measured is specific behaviour, and the list of it should be short.
- Opening: were the company name and the reason for the call stated
- Permission: was the customer asked whether they could talk
- Qualification: were the scripted questions asked
- Accuracy: did the information given match approved information
- Next step: did the call end in a specific agreement
- Record: were the outcome code and fields filled in
- Handover: was the call passed to a person where it should have been
The point of the list is that every line answers yes or no. A scale — «good, average, poor» — increases subjectivity and produces two reviewers scoring the same call differently.
There should be no more than seven or eight criteria. A long card does not get filled in, and a card that is not filled in leaves the report empty — which is the most common ending for a quality system.
Writing the scorecard
Writing the card is an hour's work, provided a few rules are followed.
- Criteria come from the scriptThe card is written after the script, not before. Requiring something the script never asked for looks unfair to the team and creates resistance.
- Separate critical from ordinarySome lines — stating unapproved information, for instance — fail a call on their own. The rest are weighted.
- Test it on sample callsTwo people score the same five calls separately. Where the results differ, the criterion is unclear — fix the card, not the people.
- Show the card to the team in advanceA hidden scoring system is discovered sooner or later and loses its credibility. The card being open is a condition of it working.
What automation can and cannot do
Automation does not replace review here, but it widens the volume that can be covered substantially.
- It can: check structure — whether the scripted steps happened in order
- It can: check whether specific statements were present
- It can: measure signals like talk time, pauses and where a call was cut off
- It can: check that the outcome code and fields were completed
- It cannot: judge whether a decision requiring context was the right one
- It cannot: identify the real reason a customer was unhappy
- It cannot: say whether departing from the script was justified
The practical model is mixed: automated checks filter every call and separate the part that needs attention; a person listens to that part. The same hour of a team lead's time then produces more, because they are listening to selected calls rather than random ones.
Building this requires calls to be recorded and processed, and that depends on the company's telephony and platform. Establish what is actually stored before designing the audit, because an audit can only be built on data that exists.
The legal side: recording, consent, retention
Recording and processing calls is processing personal data, and that is a legal decision before it is a technical one.
On the practical side three things have to be written down: which calls are recorded, how long recordings are kept, and who can access them. The system should not be built without those three answers.
A fourth question is where the data lives. Recordings and transcripts can be sensitive: a customer's financial position, their health, their personal circumstances. Access to them should not be open to the whole team.
Getting the findings back to the team
The value of a quality audit is in the change, not in the report. How the findings reach the team therefore matters.
- Individual results go to individuals, aggregate results to the team
- Each review proposes one specific change — not five
- Good examples are shared: a well-run call is the best training material there is
- A recurring mistake is not an individual problem but a script problem, and the script gets fixed
- The review runs on a steady rhythm — a once-a-month campaign of checking does not work
The fourth line is worth pausing on. If four operators out of five make the same mistake at the same point, the people are not the mistake: that part of the script is unclear, and fixing it is cheaper than training five people.
Misuse: as an instrument of punishment
The fastest way to destroy a quality system is to use it as a disciplinary tool. The result is predictable and it is the same every time.
The team learns to raise the score: scripted sentences are recited mechanically, difficult questions are not asked, risky conversations are cut short. The score rises and conversion falls — and the report does not show it, because quality scores and outcome metrics are read separately.
The protection is simple: quality scores are always read alongside outcome metrics. An operator with a high score and low conversion follows the script but does not hold a conversation — that is a coaching subject, not a disciplinary one.
Measurement
The audit itself has to be measured, or it becomes a formality within a few months.
- Coverage: what share of calls was reviewed
- Critical breaches: unapproved information, a handover that did not happen
- Breakdown by criterion: which line is failed most often
- Agreement between reviewers: how close are two scores of the same call
- The relationship between quality score and conversion — the system's main check
That last line connects the quality system to the outcome metrics. The other links in the chain are listed in telesales call metrics.
Limits
A quality audit does not change the outcome of a call — it only shows what happened. The change is made in the script, in training or in the list.
The second limit is in measurement itself. Some important things are not measurable: hesitation in a customer's voice, the tone of a conversation, an operator knowing when to stay quiet. Judging those needs a person listening, and that cannot be replaced entirely.
What Vexvon provides
On the call side Vexvon provides part of the record a quality audit needs; the rest depends on the company's telephony and recording setup.
- The scenario defines the fields to extract from a call — so «was the question asked» is answerable at field level
- The knowledge base limits what the agent may say, which structurally reduces the risk of unapproved information
- A one-sentence summary is produced from the conversation and stored on the lead record
- Outcome code, contact attempts and close reason are written
- One timeline shows the order of events
- The Excel report covers the overview, working-hours analysis and an hourly heat map
Reporting capabilities are on the analytics page, and what the agent's answers are grounded in is on the knowledge base page.
First step
Write a seven-line card and score twenty calls by hand over one week. That step comes before automation, because it is the only way to find out whether the card itself works.
Then look at which lines could be checked automatically. To discuss how that would be built on your own call flow, get in touch.