Skip to main content
Customer service training

Customer service training scorecard: criteria and weights

When a sales scorecard is copied into support, agents lose points for "not selling" and solution quality goes unseen. This guide builds a customer service training scorecard: a seven-criterion example with weights, three levels, critical criteria, chat and call versions, calibration, versioning and what the scorecard should not be used for.

October 6, 20266 min read

Short answer

A customer service training scorecard is a table of 6–8 criteria for evaluating practice conversations the same way every time. Each criterion is written as observable behaviour, has three levels (done, partly, not done), is weighted by the company, and the weights add up to 100. Some criteria are "critical" — giving wrong information or promising beyond authority, for example — and need their own discussion regardless of the overall score. Before use, the scorecard is calibrated: two assessors score the same 10 conversations and discuss the differences.

Why a sales scorecard doesn't fit support

In a sales training scorecard, most weight goes to needs discovery, objection handling and the next step. In support the key questions are different: was the problem understood, was the answer correct, did the customer leave knowing what to do, did the agent follow policy? A team that copies the sales scorecard into support ends up scoring agents on "selling", and solution quality disappears. The sales version is covered in the sales training scorecard.

It also differs from a quality form for real calls: that evaluates a conversation that has already happened, while a training scorecard is tied to the purpose of the practice scenario, with one or two criteria getting special attention in each.

Seven criteria: an example

The set below is illustrative and can be a starting point. Change the weights to fit your priorities:

  1. Clarifying the problem — 20The agent asked identify, reproduce and scope questions and found the real cause.
  2. Accuracy of the answer — 25The information given matches the knowledge base; nothing is invented.
  3. Empathy and tone — 15The problem and its impact were named; no empathy-breaking phrase was used.
  4. Authority limits — 10Promises and offers fit the authority matrix.
  5. Clarity of the solution — 15The customer knows what will happen and what they need to do.
  6. Next step and timeframe — 10A concrete timeframe was given; a follow-up was planned if needed.
  7. Channel conduct — 5Short and structured in chat; one question, one answer on a call.

The questioning criterion is detailed in clarifying questions for support agents, and the empathy criterion in empathy training.

Three levels: done, partly, not done

A five- or ten-point scale creates big gaps between two assessors: one gives a 7, the other a 5, and both think they are right. Three levels are applied more consistently: done (full points), partly (half), not done (zero). Write one example sentence per level — the assessor places the boundary by looking at the example. Some criteria may not apply in a scenario; mark them "not applicable" and leave them out of the calculation.

Critical criteria

Some mistakes must not get lost in the overall score. If an agent is strong on everything else but told the customer the wrong return policy, a score of 85 misleads. So mark two or three criteria as critical — wrong information, promises beyond authority, sharing personal data against the rules — and discuss the result separately whenever they are breached, whatever the score. This is a process rule: who discusses it, when, and when the practice is repeated.

Versions for chat and calls

The same scorecard does not work unchanged for chat and calls. On a call, not interrupting, pauses and tone of voice matter; in chat, the structure of the reply, the number of messages and delay. Keep the core criteria the same and write the channel criterion with two different definitions. That way an agent's results in the two channels can be compared.

Calibration

  1. Pick 10 conversationsAt different levels: good, average, weak.
  2. Score separatelyTwo or three assessors, without looking at each other's scores.
  3. Find the gapsWhich criteria differ by more than one level.
  4. Tighten the definitionAdd to the definition and the example sentence of the criterion with gaps.
  5. Compare with AIIf there is an AI score, compare it with the human scores on the same 10 conversations.

Versions and history

When the scorecard changes, old and new results cannot be compared directly. Give every change a version number and date, and store which version each evaluation used. A weight change is a version change too: if "accuracy of the answer" goes from 25 to 30, agents' scores can shift without their skill changing at all.

What not to use the scorecard for

A training scorecard is a development tool: it shows which skill an agent needs to practise. Making it the sole basis for hiring, pay or dismissal is wrong — a simulation is not a real customer, and both AI and human assessors can err. This is covered in detail in AI training scores and HR decisions.

An illustrative example

This is an illustrative example. A telecom company's support training used to use the sales scorecard, and agents lost points for "not offering an add-on". After switching to the seven-criterion scorecard, calibration showed two leads understood "clarity of the solution" differently. An example sentence was added to the criterion.

"Wrong tariff information" was chosen as a critical criterion. In the first month three agents breached it, and all three ran the scenario again with their lead — before facing real customers.

Common mistakes

  • Copying the sales scorecard into support.
  • General criteria like "was professional" or "was empathetic".
  • 15–20 criteria — nobody applies them consistently.
  • Launching without calibration.
  • Comparing old and new scores after a version change.

Limitations

Even the best scorecard cannot capture every nuance of a conversation: tone, humour and cultural differences do not turn into numbers. A high practice score does not mean high quality in real work. A scorecard's value lies in consistency — the same conversation should get the same score — and only regular calibration protects that.

In Vexvon AI Training

In Vexvon AI Training, the company sets its own evaluation criteria — up to 15 — and their weights; the weights come from the company, not the model. Each evaluation stores the weights and criteria version it used, so after a change it is clear which version a result belongs to. The result shows a score and explanation per criterion, mistakes tied to the agent's own sentences and an overall 0–100 score. More on AI Training.

Next step

Take your current form and ask one question of each criterion: can it be shown with a concrete sentence from the conversation? If not, rewrite it as behaviour. The quality form for real calls is covered separately in the sales and support scorecard. More articles are in the customer service training section, and we can build your scorecard together during a demo.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.