Call center scorecard template: a practical agent evaluation form
A call center scorecard is only as good as the way its criteria are written. This guide sets out six rules for a good criterion, a sample scorecard with five sections and twelve criteria, a definition card for each criterion, what to leave off the form, and a two-person test to run before launch.
The short answer
A call center scorecard is a document that sets out in advance what will be checked on a call. A good one has four parts: sections (the stages of the call), criteria in each section that ask about observable behaviour, a precise definition of each answer option, and a requirement for the evidence that supports the result. A scorecard should be about what was said and done, not about feelings.
Below is a sample form with five sections and twelve criteria. Use it as a starting point to adapt to your call type, not as something to copy as is.
Six rules for a good criterion
AWS's and Google's guidance on automated evaluation gives almost the same rules, and they hold for a human evaluator too:
- Write the criterion as a full-sentence question: not "ID validation" but "Did the agent confirm the customer's identity?"
- Make clear who is being evaluated: not "Was profanity used?" but "Did the agent use profanity?" — otherwise the customer's words count against the agent
- Ask positively: not "did the agent skip the greeting" but "did the agent greet the customer"
- Write down when each answer applies, including the "not applicable" condition
- Say whether "all" or "any" is required: if name and number must both be asked, asking only the name is not a "yes"
- Remove subjective criteria such as tone, voice or "was attentive" — they cannot be checked reliably from a transcript
A sample form: five sections
An illustrative example: a form for an inbound call about a product or service. The "yes" condition is written next to each criterion; everything else is "no", and the cases marked separately are "not applicable".
- 1. Opening1.1 Did the agent state the company's name and their own? · 1.2 Did the agent ask for or confirm the reason for the call?
- 2. Understanding the need2.1 Did the agent ask at least one question that clarified the customer's need? · 2.2 Did the agent summarise the need in their own words before making an offer? · 2.3 Did the customer have to repeat the same information? (a "yes" here is a bad result — the question is reversed; keep such criteria to a minimum)
- 3. Accuracy of the answer3.1 Did the price stated match the price in force on the date of the call? · 3.2 Did the agent explain the conditions of the price (what is included)? · 3.3 Did the agent say they would check rather than guess an answer they did not know? (no such question — not applicable)
- 4. Process compliance4.1 Where required, did the agent confirm the customer's identity? (call type where it is not required — not applicable) · 4.2 Did the agent give the mandatory information (a cancellation condition, for example)?
- 5. Closing5.1 Did the agent summarise what was agreed? · 5.2 Did the agent agree a specific next step (a date, a time or a responsible person)? · 5.3 Did the agent ask whether the customer had any other questions?
Note that criterion 3.1 cannot be checked from the transcript alone: it needs the price list in force on that date. Keep such criteria on the form, but mark that they require outside information.
A definition card for every criterion
The most skipped part of a form is the explanation of each criterion. Write a short card for each:
- Question — a full sentence, about the agent
- Why it matters — one sentence; agents will read it too
- When it is "yes" — a specific behaviour and two example phrasings (not a verbatim script)
- When it is "no"
- When it is "not applicable" — for example, the call was transferred or the customer never asked
- Non-standard cases — callbacks, transfers, the customer hanging up mid-call
- Evidence — which part of the transcript the result has to rest on
A filled-in example: one call's result
An illustrative example: a customer wants to change their internet plan. Part of the form could be filled in like this — with the transcript location that shows each result:
- 1.1 Yes — line 1: "Hello, … company, this is Nigar"
- 2.1 Yes — line 4: "What do you mostly use the internet for at the moment?"
- 3.2 Partial — line 9: the monthly price was given, but not whether the installation fee is included
- 3.3 Not applicable — the customer asked nothing the agent did not know
- 5.2 No — line 14: "Think it over and call us again" — no when, who or what for was agreed
In this format the agent looks at the line instead of arguing with the result. The team lead's conversation gets specific too: two behaviours, 3.2 and 5.2, and two lines.
What should not be on the form
- Tone of voice, intonation, "spoke with energy" — not measurable from a transcript
- "Was professional", "was attentive" — every evaluator reads them differently; break them into observable behaviours
- Whether a sale happened — that is an outcome, not a behaviour; track it as a separate measure
- Whether a CRM entry was made — not visible from the call; check it in the system record
- Any judgement about the agent's personality, character or voice
Test the form before launching it
- Pick ten callsTen calls of different types and outcomes.
- Have two people fill it in independentlyWithout seeing each other's answers.
- Count the differences by criterionOn which criterion did the two disagree most?
- Fix the definition cardThe cause is usually in the card: the "not applicable" condition is missing, or "yes" is described too broadly.
- RepeatUntil the differences shrink. Only then is the form announced to agents.
Limits
The sample form is not universal: sales, support, returns and complaint calls need different criteria. Apply one form to every call and "not applicable" statuses will multiply, leaving the score based on few criteria. Criterion weights and the effect of critical errors are a separate decision — make it once the form is ready. A form's result should never be the only basis for a decision about an agent.
Moving the form into Vexvon Audio Analyzer
In Vexvon Audio Analyzer the form is entered not as separate fields but as a standard the company writes in free text; the system breaks it into steps. So write the definition cards straight into the text — for example: "Next step: the agent agrees a specific date, time or responsible person. If the customer hung up mid-call — not applicable."
- Each step gets a "met", "partial", "missed" or "not applicable" status
- The status comes with a comment and references to transcript lines — the reviewer looks at the evidence
- Changing the standard creates a new version; each call keeps the text it was evaluated against
On the five layers of an evaluation result, see what AI call analysis is.
First step
Pick your most common call type and write at most twelve criteria, starting from the five sections above. Fill in a definition card for each and run a two-person test on ten calls. More evaluation pieces are in the agent scoring section; to try the form on your own calls, get in touch.
Further reading on this topic: weighted call center scorecard, not applicable in call QA scoring, sales vs service call scorecard.