"No", "Not applicable" and "Unable to assess" in call QA scoring
Marking "no" for objection handling on a call without an objection penalises the agent; counting a recording gap as "not applicable" hides the problem. This guide explains how not applicable works in call QA scoring alongside "no" and "unable to assess": definitions, the effect on score and coverage, a decision tree, an illustrative calculation and how "not applicable" gets misused.
The short answer
Three statuses have to be kept apart in an evaluation. "No" — the criterion applies to this call, there is evidence, and the agent did not do what was required; it lowers the score. "Not applicable" — the criterion's condition never arose on this call (the customer raised no objection, the call was transferred); the criterion leaves the calculation and this is not a gap. "Unable to assess" — the criterion may apply, but the recording or transcript is not good enough to decide; the criterion leaves the calculation, but it reduces evaluation coverage and is counted separately with its reason.
A form that mixes these three either penalises agents for mistakes they did not make or hides real problems under "not applicable".
How mixing them does damage
Two illustrative cases. First: a customer asked about a plan, agreed and ended the call. The form has a criterion "did the agent handle the objection correctly", and it gets "no" — because no objection was handled. But there was no objection. The agent loses points for a mistake that did not happen.
Second: there is a ten-second break in the middle of the recording, and it falls exactly where the mandatory disclosure should have been made. If the criterion gets "no", the agent is penalised for a technical problem. If it gets "not applicable", nobody notices whether the disclosure was made. The right status is the third one — "unable to assess" — and the call should go to human review.
Three statuses: definition and effect
- NoThe condition arose, there is evidence, the behaviour is missing. Score: 0 points, stays in the denominator. Coverage: counts as assessed. A topic for the conversation with the agent.
- Not applicableThe criterion's condition did not arise on the call. Score: leaves the numerator and denominator. Coverage: no effect, because there was nothing to assess. The condition is written in the criterion's definition card in terms of the customer's behaviour or the call type.
- Unable to assessThe condition may have arisen, but the evidence is unreliable: audio quality, a break, an unidentified speaker, language. Score: leaves the numerator and denominator. Coverage: reduced. The reason is recorded and the call goes to human review.
A decision tree for choosing the status
- 1. Did the criterion's condition arise on this call?No — "not applicable". Yes or unknown — go to the second question.
- 2. Is that part of the call reliable?Inaudible, cut off, speaker unclear — "unable to assess". If reliable — go to the third question.
- 3. Did the required behaviour happen?Fully — "yes". Partly (the name was asked, the number was not) — "partial", if the form allows it. Not at all — "no".
An illustrative calculation
An equally weighted ten-criterion form. The call's result: six "yes", one "no", two "not applicable", one "unable to assess".
- Criteria that apply10 − 2 = 8
- Criteria assessed8 − 1 = 7
- Score6 ÷ 7 ≈ 85.7%
- Evaluation coverage7 ÷ 8 = 87.5%
Had "not applicable" and "unable to assess" also been counted as "no", the score would be 60% — the agent would unfairly lose points on three criteria. Had "unable to assess" been silently recorded as "not applicable", the score would be the same, but nowhere would it show that one criterion was never checked.
Misusing "not applicable"
"Not applicable" is most often misused in two ways. First, a step the agent skipped is recorded as "not applicable": the customer objected to the price, the agent changed the subject, and the criterion was marked "no objection". Second, the form does not fit the call type, and half the criteria come back "not applicable" on every call — leaving the score based on very few criteria.
Both are detected the same way: track the share of "not applicable" per criterion. If it suddenly rises on one criterion, or is much higher for one agent than for the others, a person should check those calls.
The AI "not applicable" trap
AWS's documentation on generative-AI evaluations describes an important behaviour: when the model cannot determine which of the given answer options fits, it selects "not applicable", even if the question is mandatory; the question is then excluded from scoring and its weight is redistributed to the others. In other words, the system's uncertainty can quietly turn into a higher score. The same documentation recommends stating explicitly in the instructions when each answer option is chosen. That describes one specific product, but the risk is general: track the share of "not applicable" separately in any system.
An "unable to assess" report
- Count by reason: audio quality, a break, an unidentified speaker, a language the form was not designed for
- Split by shift, line or call source — to find a technical problem quickly
- No negative effect on the agent's score; shown in a separate column in the agent's average
- Every such call goes to human review, at least on the result's critical criteria
- If the share rises, check the recording and telephony settings first, not the agents
Limits
Three statuses make a form more complex and need definition cards and a test run so evaluators read them the same way. If a "partial" status is added, its point value and when it is given have to be written down. And no status rule makes a decision about an agent automatic: the result has to be checked by a person, and the agent must be able to appeal.
Statuses in Vexvon Audio Analyzer
In Vexvon Audio Analyzer every step of the standard gets one of four statuses: "met", "partial", "missed" and "not applicable". "Not applicable" never counts against the agent. There is no separate "unable to assess" status — instead there are signals: poorly recognised lines are marked in the transcript, and a line whose speaker could not be determined stays "unknown". Put such calls in the review queue and have a person confirm the result.
How statuses affect a weighted calculation is covered in the weighted scorecard guide; setting aside calls unfit for analysis in call center QA sampling.
First step
For every criterion on your form, write the "not applicable" condition in one sentence in terms of the customer's behaviour or the call type. Then calculate the share of "not applicable" per criterion in last month's results — fix the definition card starting with the criterion that has the highest share. More in the agent scoring section; get in touch.
Further reading on this topic: call recording quality speech analytics, evidence-based call scoring.