Skip to main content
Support quality & coaching

Measuring coaching effectiveness: before and after

"After coaching the average score rose from 62 to 78" proves nothing on its own: the later period may simply have had easier calls than the earlier one. This article covers measuring coaching on the target behaviour, comparable groups of calls, an illustrative easy-call-trap example with checked figures, the role of form versions, choosing periods, a comparison group, four reasons why coaching may show no result, and how to write up the outcome honestly for management.

September 29, 20267 min read

Short answer

To measure whether coaching worked, compare calls from before and after the coaching using the same criteria, the same version of the evaluation form, and comparable groups of calls. The most common mistake is that the later period happens to contain easier calls: the score rises, but the cause is the call mix, not the coaching.

In other words, "after coaching the average score rose from 62 to 78" proves nothing on its own. On which criterion, which call type, how many calls and which form version — until those questions are answered, the result may be chance or measurement error.

What to measure: the target behaviour, not the score

A coaching plan targets one or two specific behaviours. The measurement should target those behaviours too, not the overall score. The overall score adds up ten criteria, and random movement in the others can hide the change in the target behaviour.

  1. Target criterionThe behaviour in the plan: for example, "the need was clarified before the price was given". The main measure is the share of calls where this criterion was met.
  2. Applicable callsThe denominator is only the calls where the criterion applies. On a call with no price question, the criterion is not applicable and is left out of the calculation.
  3. Control criteriaOne or two criteria the coaching did not touch are tracked as well. If they rose just as much, the cause may be a general change rather than the coaching.
  4. Overall scoreShown for context, but not presented as the result of the coaching.

Comparable groups of calls

  • Same call type: do not compare sales calls with support calls
  • Similar difficulty: complaints and simple enquiries are counted separately
  • Same language: if one period has many Russian-language calls and the other few, the result is distorted
  • Similar recording quality: the share of "unable to assess" calls is shown separately for both periods
  • Enough calls: the number of calls where the target criterion applied is stated for each period; a percentage from a handful of calls is noise

Illustrative example: the easy-call trap

The figures below are an illustrative example, not a real agent's data. Target criterion: "the need was clarified before the price was given". Each period has 20 calls.

  1. Before coaching12 complex calls, criterion met on 3. 8 simple calls, 2 of which had no price question (not applicable); met on 3 of the remaining 6. Total: 6 / 18 = 33%.
  2. After coaching4 complex calls, met on 2. 16 simple calls, 4 not applicable; met on 11 of the remaining 12. Total: 13 / 16 = 81%.
33% → 81%Overall share: coaching looks like a big success
25% → 50%Complex calls only
64%The after result, re-weighted to the before period's call mix

The improvement is real — the share rose on both call types. But part of the jump from 33% to 81% comes from complex calls falling from 12 to 4 in the second period. Re-weighting the after result to the before mix — (12 complex × 50% + 6 simple × 92%) / 18 — gives about 64%. That is the figure that belongs in the report, not 81%.

When the form changes, results do not compare

The evaluation criteria may change at the same time as the coaching: a criterion is clarified, a step is added, the definition of an answer option changes. The before and after scores were then measured with different rulers.

  • Record which form version each evaluation used
  • The form should not change during the period in which coaching is measured; if it must, re-evaluate the earlier calls with the new version
  • AWS's evaluation-form documentation also notes that historical evaluations completed with a different scoring mode cannot be compared directly

Choosing periods and waiting time

  1. Before periodTwo to four weeks immediately before the coaching. Going further back pulls in other changes — season, a new product, new pricing.
  2. TransitionThe first few days after coaching are usually excluded: the agent is only starting to try the new behaviour and results are unstable.
  3. After periodAfter the transition, the same length as the before period. If possible, a second measurement a month later — has the behaviour become a habit, or slipped back?

A comparison group

The most reliable check is a comparison with similar agents who did not get the coaching. If their target criterion improved just as much over the same period, the cause is not the coaching — perhaps the customer mix changed or a new campaign started.

In a small team that is not always possible. Then at least look at the control criteria and note other changes — a new script, a new product, a shift change — in the report.

If there is no result: four causes

Sometimes proper measurement shows "no change". That does not mean the coaching was wasted — finding the cause improves the next plan.

  1. The behaviour was written too broadly"Be more attentive to the customer" cannot be measured or practised. Rewrite the behaviour as one observable sentence.
  2. Practice did not resemble real callsThe role play used an easy customer, while real calls bring customers who push back. Match the practice to the hardest situation that comes up most often.
  3. Conditions did not changeThe agent knows the new behaviour, but the system gives no time for it or the script asks for something else. That is a process issue, not a coaching one.
  4. The measure is not sensitiveThe criterion's definition is so broad that both good and bad answers come out as met. Tighten the criterion and re-evaluate the earlier calls with the new version too.

How to write up the result

  • The target criterion and its definition
  • The number of applicable calls and of "unable to assess" calls in each period
  • The result split by call type
  • The form version, and whether it changed during the period
  • Other changes that happened in the same period
  • Conclusion: "improved", "no change" or "inconclusive" — the last is an honest and useful answer too

Limits

  • A change measured on few calls may be chance; sharp percentages deserve extra caution
  • Higher sales do not prove coaching worked on their own — sales also depend on product, price and lead quality
  • AI evaluation has its own error; a person should check a few calls on the target criterion too
  • The result of coaching should not be the sole basis for an agent's reward or discipline

What Vexvon Audio Analyzer offers

  • On every call, each step of your call standard is judged as met, partial, missed or not applicable — the status needed to get the target criterion's denominator right is there
  • Call standards are versioned: every call keeps the text and version it was judged by, and scores are compared only within one version
  • When the standard changes, the existing transcripts of older calls can be re-analysed against the new one — without transcribing the audio again
  • Each judgement comes with transcript lines, which makes a human spot check easier

More: Vexvon Audio Analyzer. How to write the coaching plan itself is covered in call center coaching plan.

First step

Before the next coaching session, write down the target criterion, the call type to compare and the two periods. Afterwards, show the result by call type rather than as one overall share. This article belongs to the support quality and coaching section. To try it on your own recordings, get in touch.

Further reading on this topic: fair agent performance comparison, QA scorecard versioning.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.