Weighted call center scorecard: how to choose weights and calculate
Skipping a greeting and quoting a wrong price do not have the same effect, yet an equally weighted form counts them the same. This guide builds a weighted call center scorecard: choosing weights, an exact scoring rule, removing "not applicable" criteria from the denominator, evaluation coverage, a step-by-step calculation and edge cases.
The short answer
In a weighted call center scorecard, each criterion's or section's effect on the overall score is set by its importance to the customer, the business and risk. Skipping the greeting and quoting a customer the wrong price do not have the same effect, yet an equally weighted form counts them the same. Choosing weights is only half the job: you also have to write down in advance the denominator (what the score is divided by), how "not applicable" criteria are removed and how evaluation coverage is shown.
Why equal weights mislead
Take a ten-criterion form where each criterion counts for 10%, and two illustrative calls. In call C the agent skipped the greeting and the introduction; everything else was right. In call D everything was right except that the customer was quoted an old price. With equal weights C gets 8/10 = 80% and D gets 9/10 = 90%. The call with the wrong price looks better.
If a manager reads that ranking and talks to the agent from call C, they will miss the real risk — a wrong commitment made to a customer. Weights exist to correct exactly this distortion.
How to choose weights
- Rank the impactFor each section ask: if this is not done, what is lost for the customer, the revenue or legal risk? Rank in words first, without numbers.
- Use few levelsThree or four weight levels are enough (say 10, 20, 25, 35). Figures like 37% or 13% look precise but cannot be justified.
- Make them total 100Section weights add up to 100; within a section, criteria can share equally.
- Test with scenariosCalculate a few calls you already consider "bad" and "good". If the result contradicts your judgement, check the criterion before the weight.
- Record the versionA change of weights is a new form version; it is not compared directly with old scores.
An illustrative form and scoring rule
The sample form has five sections with two criteria each. Weights: opening 10, understanding the need 25, accuracy of the answer 35, process compliance 20, closing 10 — a total of 100.
- For each criterion: "yes" = 1, "partial" = 0.5, "no" = 0
- A section's result = the average of its criteria × the section weight
- A "not applicable" criterion leaves both the numerator and the denominator; if a whole section does not apply, its weight leaves the denominator too
- Final score = weight earned ÷ weight assessed × 100
- The weight assessed is shown separately — that is the evaluation coverage
A step-by-step calculation
Call A (every criterion applies): opening 2/2 yes; need — one yes, one no; accuracy — one yes, one partial; process 2/2 yes; closing — one yes, one no.
- Opening(1 + 1) ÷ 2 = 1.0 × 10 = 10
- Need(1 + 0) ÷ 2 = 0.5 × 25 = 12.5
- Accuracy(1 + 0.5) ÷ 2 = 0.75 × 35 = 26.25
- Process(1 + 1) ÷ 2 = 1.0 × 20 = 20
- Closing(1 + 0) ÷ 2 = 0.5 × 10 = 5
- Total10 + 12.5 + 26.25 + 20 + 5 = 73.75 — denominator 100, coverage 100%
Call B is identical, but on this call type identity confirmation and the mandatory disclosure are not required — the process section is "not applicable". Weight earned is 10 + 12.5 + 26.25 + 5 = 53.75, weight assessed is 80. Score: 53.75 ÷ 80 × 100 ≈ 67.2.
Weighted, call C gets 90 and call D gets 82.5 — the ranking is corrected. But D still looks high. That shows the limit of weighting: a serious breach cannot be caught with weights alone; it needs a separate critical-error rule.
Show section results, not just the total
A weighted total is convenient for comparison, but it teaches the agent nothing. Call A's 73.75 says "average"; the section results say something else: opening and process complete, accuracy 75%, need and closing 50%. The conversation with the agent is built on those two weak sections.
So keep three layers in the report: the total (for ranking), section percentages (for the coaching topic) and criterion statuses with evidence (for the conversation itself). The same split helps with an agent's monthly average: it may be 80, but if "accuracy of the answer" is consistently below 60, that is the most important thing the weighting hides.
Edge cases
- Every criterion is "not applicable" — no score is calculated; the call is recorded as "not evaluated", not as 0 or 100
- Coverage is below the company's threshold — the score is shown, but flagged "limited coverage" and left out of the agent's average
- Evidence is missing (audio quality, speaker not identified) — the status is "unable to assess", not "no"; it leaves the score calculation and is counted separately
- A reversed question ("did the customer have to repeat themselves") — flipped before scoring, so "yes" is always the good result
- A critical error — handled by a separate rule, regardless of the weighted calculation
Limits
Weights are a decision, not a truth. Copying another company's weights or a vendor's sample formula will not reflect your risk. As soon as weights change, scores are no longer directly comparable with the previous period. And no weighting scheme turns a score into the only basis for a decision about an agent — it only shows which call and criterion to look at.
The score in Vexvon Audio Analyzer
Vexvon Audio Analyzer gives every step a "met", "partial", "missed" or "not applicable" status and calculates the 0–100 score from those statuses; "not applicable" never counts against the agent. Assigning separate weights to criteria is not a feature of the product today — if you want a weighted calculation like the one above, compute it from the step statuses in your own report.
For the content of the form, see the call center scorecard template; for how criteria are written, what AI call analysis is.
First step
Rank your form's sections by risk in words, choose three or four weight levels and score five calls from last month both equally and weighted. If the ranking changes, the weights are doing their job. More in the agent scoring section; to build the calculation together, get in touch.
Further reading on this topic: not applicable in call QA scoring, critical errors in call center QA.