Critical errors in call center QA: how they should affect the score
A high score can hide a serious breach, and an automatic flag is not yet a breach. This guide covers what a critical error is and is not in call center QA, three models for its effect on the score (automatic fail, cap, separate measure), the flow from flag to confirmation, an escalation matrix and how to write a critical criterion.
The short answer
A critical error is a breach the company has approved in advance that can seriously harm the customer, the company or compliance with the law: an unapproved promise, disclosing confidential information to a third party, skipping a mandatory disclosure. Adding it to a weighted score as an ordinary criterion is not enough — a high score can hide a serious breach. But an automatic flag is not a confirmed breach either: what the system says "may have happened" should reach an agent's record only after a person has checked the evidence.
A practical rule: track critical errors as a separate measure, choose the rule for their effect on the score in advance, and put every flag through a confirmation step before it becomes a result.
What a critical error is and is not
The list should be short — if everything is critical, nothing is. It usually includes:
- An unapproved promise or commitment: a discount that does not exist, a price no longer in force, a deadline that cannot be met
- Disclosing confidential or personal information to someone whose identity was not confirmed
- Skipping a disclosure or step required by law or internal rules
- Insults, threats or discriminatory remarks towards the customer
- Presenting an operation as "agreed" despite the customer's explicit objection
Not critical: a missed greeting, a long call, a tone that seems "cold". These are quality problems, not breaches, and ordinary criteria handle them.
Three models for the effect on the score
- Automatic failA confirmed critical error sets the score of the section or the whole form to zero. AWS evaluation forms offer this as "automatic fail". The advantage is a clear signal. The risk is that one wrong flag zeroes the whole call, so it should never be applied without confirmation.
- CapA call with a critical error cannot score above a set limit (say 50). The rest of the call stays visible, but a high score does not hide the breach.
- Separate measureThe score does not change; critical errors are counted separately and shown on their own line in the report: "share of calls without critical errors". The quality of the agent's behaviour and risk events are not mixed.
For many companies the clearest combination is a separate measure plus a cap or automatic fail applied after confirmation. Whatever you choose, it has to be written down and the same for every agent.
From flag to confirmation
- FlagThe system or an evaluator records a "no" or a suspicion on a critical criterion. Nothing is written to the agent's record at this stage.
- Checking the evidenceAn assigned person looks at the transcript lines and, where needed, the recording: were the speakers identified correctly, was the sentence taken out of context, was the price really different on that date.
- Decision"Confirmed", "not confirmed" or "unable to assess". The reason is recorded.
- ConsequenceOnly when confirmed does the scoring rule apply and the escalation matrix start.
- Telling the agentThe agent sees the evidence and can appeal.
An escalation matrix
An illustrative matrix — who does what, and when, after a confirmed critical error:
- Unapproved promise or wrong price — the team lead arranges a correction call to the customer within one working day; the price owner is informed
- Disclosure of confidential information — immediately to information security and legal; the customer is notified according to company rules
- Missed mandatory disclosure — compliance decides how the information will reach the customer
- Insults or threats — team lead and HR; an apology call to the customer
- Every case — a root-cause note: training, script, knowledge base or a deliberate breach
Three models on one call
Recall call D from the weighted scorecard guide: everything was right, but the customer was quoted an old price, and the weighted score is 82.5. Suppose a wrong price is on the critical-error list and a person has confirmed it.
All three models lead the conversation with the agent to the same place — the wrong price — but they look different in the report. Automatic fail pulls the average down sharply; a separate measure keeps the quality of behaviour apart from risk. Whichever you pick, keep the same rule for the whole period.
How to write a critical criterion
AWS's guidance on automated evaluation recommends giving automatic-fail questions the most attention, because one answer changes the score of the whole form. In practice that looks like this:
- Describe the breach as a behaviour: "the agent offered a discount that is not on the price list"
- If outside information is required, say so: the price list, campaign conditions, identity rules
- Add non-standard cases: what happens if the agent corrected the mistake later in the same call?
- Write the "not applicable" condition: for example, a call type where identity confirmation is not required
Limits
Transcript-based checks can raise false flags: speakers mixed up, a sentence recognised only in part, an agent repeating the customer's words. So a critical error must never become a consequence without confirmation. Which breaches carry legal consequences depends on the country and sector — have the list approved by a lawyer. An AI result is not a legal verdict and should not be the only basis for a disciplinary decision about an agent.
Critical steps in Vexvon Audio Analyzer
In Vexvon Audio Analyzer every step of the standard gets a status, a comment and references to transcript lines — so a "missed" status on a critical step reaches the reviewer together with its evidence. There is no separate automatic rule for critical errors (automatic fail, a cap) in the product: name the critical steps clearly in the standard's text and put "missed on this step" into your review-queue rule, so a person confirms the result.
The limit of weighting is explained in the weighted scorecard guide, and the difference between statuses in no, not applicable and unable to assess.
First step
Write your critical-error list in no more than five items, choose one of the three models and set a review deadline. Then check by hand the calls from last month that got "no" on a critical criterion: how many were confirmed? That ratio tells you how far to rely on the flags. More in the agent scoring section; get in touch.
Further reading on this topic: call center QA appeals, financial services call quality monitoring.