Skip to main content
Call QA strategy

Automated call quality assurance vs manual review: who does what

Automated call quality assurance cuts listening time, but it should not take the decision away from people. This guide separates what people and AI each do well, and sets out a seven-line division-of-labour matrix, queue rules that pick which calls go to a person, and the two most common wrong splits.

September 29, 20267 min read

The short answer

Automated QA (quality assurance) takes over the first read of each call and the check against observable criteria. The decisions stay with people: which result is correct, what to discuss with the agent, whether the criterion itself needs to change. With the right split, AI cuts listening time and the quality specialist spends that time on harder work — verification, coaching and improving the standard.

A wrong split takes one of two forms: writing the AI score straight into the agent's performance record, or distrusting the AI and listening to every call by hand anyway. The first is unfair; the second is pointless.

Why human listening is valuable

When an experienced quality specialist listens to a call, they do more than tick criteria. They hear the customer's real problem, why the agent chose the route they did, and whether the rule itself fits the situation. A new type of problem — a campaign condition customers keep misunderstanding, say — is usually not on the form yet, and it is a person who finds it.

The human weakness is volume and consistency. There is a limit to how many calls someone can listen to in a day, fatigue changes judgement, and two specialists can score the same call differently. That is why manual review is built on a sample, and the sample is rarely random — calls with complaints get picked.

Where AI is strong and where it is weak

AI is strong on precisely written criteria that are visible in the transcript: was there a greeting, were the price conditions stated, was a next step agreed. It asks every call the same questions and does not get tired. Its weak spots are known too: it cannot read tone from text, it does not know information that is not in the transcript (a CRM record, a payment status), it reaches wrong conclusions when speakers are mixed up, and it is inconsistent on subjective criteria such as "was the agent attentive".

AWS's documentation on generative-AI evaluations recommends the same split: it notes that AI-filled evaluations are not fully accurate, advises reviewing a sample before acting on the results, and recommends keeping a manual evaluation process to catch drift over time.

A division-of-labour matrix

The split below is a starting point for most teams. Adapt it to your process, but put a name on every line:

  1. Writing the standard — peopleThe quality manager writes the criteria, the answer options and the "not applicable" conditions. AI does not invent the standard.
  2. First pass — AIFor every analysed call, a status, a comment and a reference to a transcript line are prepared for each criterion.
  3. Review queue — a rule, then a personA rule written in advance picks which calls go to a person (see below). The specialist opens the queued call.
  4. Confirming or correcting — peopleThe specialist accepts or changes the AI status and records why.
  5. Talking to the agent — team leadCoaching works from the confirmed result: a specific behaviour and an example call.
  6. Appeal — a second personIf the agent disputes the result, someone other than the first evaluator reviews it.
  7. Changing the standard — peopleRecurring disagreements lead to rewriting the criterion; the new version is not mixed with the old one.

Which calls should go to a person

People's time is limited, so the rule for entering the queue has to be written down. Example rules:

  • A step treated as critical came back "missed" (a mandatory disclosure, for example)
  • The score is below a threshold the company has set
  • The transcript has many "unknown" speakers or poorly recognised lines — the result may be unreliable
  • Most criteria came back "not applicable" — either the call type was chosen wrongly or the form does not fit
  • A call linked to a customer complaint
  • A small random control sample every week — including high-scoring calls

The last rule matters. If you only look at low-scoring calls, you will never see the steps the AI marked as "met" that were not actually met. A random sample catches the system's errors in both directions.

Two wrong splits

The first mistake is "AI does everything". The score goes automatically into the agent's monthly figures, nobody checks it, and the agent never sees the evidence. Within a few weeks agents learn to talk to the form — whether or not that helps the customer — and trust is gone.

The second mistake is "we don't trust AI". The system is set up, but the QA team still listens to every call from the start and treats the AI result as an interesting extra. Costs go up and no time is saved. The way out is to measure trust: on which criteria AI and people agree and on which they do not — and to reduce human work only where they agree.

The rhythm of a week: an illustrative example

The example below does not describe a real company — it is built to show how the split looks in daily work. Picture a small support team with one quality specialist and three team leads.

  1. MondayLast week's recordings are uploaded and analysed. The specialist does not read the results yet; they list the calls the queue rule has picked.
  2. Tuesday – WednesdayThe queued calls are opened. In each one the AI statuses are checked against the evidence lines, and any correction is recorded with a reason. The random control sample is checked on these days too.
  3. ThursdayTeam leads hold short sessions with agents based on the confirmed results: one behaviour, one example call, one exercise.
  4. FridayThe specialist reviews the week's corrections: on which criterion did AI and people disagree most? That criterion is a candidate for rewriting next week.

In this rhythm most of the human time goes to verification and conversation rather than listening. The specialist's correction notes create value of their own: after a few weeks they are the best source for knowing which criteria work reliably.

Limits

This split does not replace human decisions; it focuses them. An AI score should never be the only basis for pay, discipline or dismissal. A result produced from a poor recording should not be used against the agent. A person changing a result is normal and should be recorded — those records later show how clearly the criteria were written.

This split in Vexvon Audio Analyzer

Vexvon Audio Analyzer does the first-pass part of the split: it turns an uploaded call into a transcript, separates agent and customer, and gives each step of the company's standard a status, a comment and a reference to a transcript line. Poorly recognised lines and uncertain speakers are marked — useful signals for a queue rule. After a criterion changes, an existing transcript can be re-evaluated without going back to the audio.

The queue rule, the appeal and the conversation with the agent are your process. For the difference between the five layers, see what AI call analysis is.

First step

Fill in the matrix with your team and put a name on every line. Then, for two weeks, compare the AI result with a human evaluation on the same calls. Reduce human work on the criteria where agreement is high, and rewrite the ones where it is weak. The overall structure of an audit is in our guide to AI call quality audits; to test it on your own calls, get in touch.

Further reading on this topic: call QA coverage vs accuracy, AI call scoring calibration.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.