Skip to main content
Agent scoring

Fair agent performance comparison: few calls vs many calls

One call can move a 12-call average by 4 points and a 140-call average by 0.3. This guide shows the role of sample size, call type, shift and language in a fair agent performance comparison, six rules, bands instead of ranks, and an illustrative line from an agent report.

September 29, 20266 min read

The short answer

For a fair agent performance comparison, the average score on its own is not enough. Four things have to be shown alongside it: how many fit calls the score rests on, the type and complexity of the calls, the shift and line, and the language. For an agent with few calls, one bad or one good call moves the average by several points; for an agent with many calls it changes almost nothing. So use a minimum sample threshold, compare within the same call type, and use bands instead of ranks.

How much one call moves the average

An illustrative example. Agent A has 12 evaluated calls in a month, averaging 88. Agent B has 140 calls, averaging 82. A ranks higher. Now suppose one call of each scored 40 where their usual level would have been.

  1. Agent A(11 × 88 + 40) ÷ 12 = 1,008 ÷ 12 = 84. One call pulls the average down by 4 points.
  2. Agent BOne call scoring 40 instead of 82 changes the average by 42 ÷ 140 = 0.3 points.
4 pointseffect of one call: a 12-call average
0.3 pointseffect of one call: a 140-call average
88 → 84A's average after one call

So the gap between A's 88 and B's 82 can vanish with a single call of A's. An agent with a small sample ranking high or low often reflects chance rather than their work. And a small-sample agent at the top of this month's ranking will most likely drift back towards the middle next month — not a drop in performance, just the nature of small samples.

Call type and complexity

Even with the same number of calls, two agents' calls are not the same. One agent's line gets mostly short "what are your opening hours" queries; another's gets complaints and payment problems. On complex calls criteria are harder to meet and averages are naturally lower. A ranking that ignores this penalises the agent who takes the hardest calls.

The fix: compare within call type. Compare agents with each other on "complaint" calls, and with each other on "information request" calls. If an agent has only two calls of a type, their result for that type is not shown.

Shift, line and language

  • Shift: the night line may get more noisy mobile calls — the share of calls fit for analysis is lower
  • Line and queue: a VIP line, a returns line and a sales line are evaluated with different criterion modules
  • Language: transcript quality can depend on language; results for calls in a different language are shown separately
  • Experience: an agent in their first three months is compared with their own previous month, not with an experienced colleague

Rules for a fair comparison

  1. Minimum sampleAn agent with fewer fit calls than the company's threshold is left out of the ranking; their result shows as "not enough data".
  2. n next to every scoreThe number of calls an average rests on is always written next to it.
  3. Within the same typeComparison is done in groups by call type, line and language.
  4. Bands, not ranksInstead of 1st, 2nd, 3rd, three bands: "above expected", "as expected", "needs attention". Small differences look big in a ranking and disappear in a band.
  5. A longer periodFor agents with few calls, use a quarter instead of a month.
  6. Against their own historyAn agent's development is compared with their own previous period, not with other people.

A sample report line

An illustrative agent line for one month: "Nigar — support line, day shift. Fit calls: 96. Base score: 84 (band: as expected). Support module: 71 (band: needs attention — the 'confirming the solution' criterion). Complaint calls: 11, not shown separately (below the minimum). Last month: base 81, module 68."

That line tells a team lead three things: the agent is at the expected level overall, the weak spot is one specific criterion, and there is progress on last month. Nowhere is there a direct rank comparison with another agent.

Side effects of a public ranking

A ranking shown to the whole team creates competition, but not always in the direction you want. Agents may start avoiding hard calls, transferring difficult customers quickly, or "meeting" the form's criteria even when it does the customer no good. Small-sample agents end up embarrassed or praised because of one random month.

A practical approach: individual results are seen by the agent and their team lead; at team level, only the split across bands and shared weak criteria are shared. If you want recognition, show a specific good call as an example — not a ranking position.

Limits

There is no universal number for the minimum sample: it depends on your call volume, the number of criteria and what the result is used for. Bands are not definitive for agents near a boundary either. And no comparison should be the only basis for a decision about an agent — it only shows which agent to talk to, and about what.

Vexvon Audio Analyzer and agent comparison

In Vexvon Audio Analyzer the score is per call: each call gets statuses on the standard's steps, evidence and a 0–100 score, and for fitness there are the markers for poorly recognised lines and "unknown" speakers. Per-agent rankings and averages are not built in the product — grouping calls by agent, shift and call type and applying the rules above is the job of your report. Because every call keeps which version of which standard it was evaluated against, groups can be split by version as well.

Reliable coverage at agent level is in call center QA sampling, and modules by call type in the sales vs service scorecard. For conversion metrics, see our sales call metrics guide.

First step

Add two columns to your current agent ranking: the number of fit calls and the main call type. Then look: how many of those at the top and bottom have a small sample or a different call flow? More in the agent scoring section; get in touch.

Further reading on this topic: QA scorecard versioning.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.