Skip to main content
Reliability & pilots

Which metrics measure an AI voice agent's work

A high resolution rate can mean the agent is keeping difficult calls to itself. This guide explains why a single indicator misleads, the definition and calculation of five indicators, splitting escalation by reason, a weekly panel, an illustrative two-week comparison and the limits of measurement.

September 30, 20266 min read

The short answer

An AI voice agent's work cannot be measured with one number. Five indicators must be read together: the share of calls the agent finished without handing over (resolution), the share routed to the right place (routing accuracy), the share and reasons of calls handed to a person (escalation), callers hanging up while talking to the agent (abandonment), and wrong answers found in sample checks (observed errors).

The key rule: a high "resolution" rate is not good in itself. If the agent also keeps difficult calls to itself, resolution goes up — and so does customer dissatisfaction. This article gives the definition, calculation and limits of each indicator.

Why a single indicator misleads

Each indicator pulls in one direction. To raise resolution, the agent can hand over less — but wrong answers increase. To reduce errors, the agent can hand over everything — but resolution falls and operators are overloaded. To reduce hang-ups, the agent can drag the call out. So indicators are read in pairs: resolution with observed errors, escalation with routing accuracy.

1. Resolution rate

Definition: the share of calls the agent finished without handing over, where the caller's need was met. The second part matters: a call simply ending is not resolution. There are two ways to check: the "outcome" field extracted from the call (for example "information given", "booking taken"), and a repeat call from the same number within a short time — a repeat call often shows the first was not resolved.

2. Routing accuracy

Definition: the share of calls handed over by the agent that went to the right department or person. This indicator is not measured automatically — it is calculated by sample checks: each week a set number of handed-over calls are listened to or read, and whether each was correct is recorded. Operators' "this call wasn't for me" signals are an additional source.

3. Escalation rate and reason

Definition: the share of calls handed to a person. But the reasons matter more than the number. Separate them: the caller asked for a person, the topic is out of scope (planned), the agent did not understand (a problem), a technical error (a problem). Planned escalation is not a failure — it shows the agent working correctly.

4. Hang-ups while talking to the agent

Definition: the share of cases where the caller ended the call before reaching their goal. Separate this from ordinary "queue abandonment" — the agent answers immediately, so hang-ups here come not from waiting but from the conversation itself: a long introduction, not being understood, going round in circles. Look at when the hang-up happens: in the first 10 seconds, after a repeated question, while waiting for a handoff.

5. Observed errors

Definition: the share of calls in a sample check where the agent gave wrong information, broke its boundary or did not hand over correctly. The word "observed" matters: this is not the number of all errors but those found in the sample reviewed. If the sample is small the figure is imprecise — so every error is analysed individually and the percentage is read as a trend.

The weekly indicator panel

  1. VolumeThe number of calls reaching the agent — for context.
  2. ResolutionRate and trend, alongside the repeat-call ratio.
  3. EscalationRate and the split by the four reasons.
  4. AbandonmentRate and at which moment.
  5. ErrorsCalls reviewed, errors found, a short note on each.

Keep the panel on one page. A manager reading it should answer three questions in five minutes: is the agent doing more work, is it making fewer mistakes, and do callers keep talking to it?

Illustrative example: comparing two weeks

Not real customer figures. Week one: 1,000 calls reached the agent, resolution 55%, escalation 38% (half of it "did not understand"), abandonment 7%, 3 errors in 50 calls reviewed. After fixing the knowledge base and scenario, week two: resolution 60%, escalation 34% (a quarter of it "did not understand"), abandonment 6%, 1 error in 50 calls reviewed.

55% → 60%resolution rate
50% → 25%"did not understand" share of escalations
3 → 1observed errors per 50 calls

The main change is not in resolution but in the reason for escalation: the agent "does not understand" less often. That shows the fix worked.

How to run the sample check

The sample check must be random, not made of "interesting" calls. Each week pick 30–50 random calls from the call list and answer three questions for each: was the information correct, was the boundary kept, and if a handoff was needed, did it happen? Also review every call with a complaint or an operator signal separately, but do not mix them into the random sample — otherwise the error rate is artificially inflated.

Measuring separately by call type

Resolution may be 90% on opening-hours questions and 50% on price questions. The overall figure shows 70% and gives a false picture of both. If the scenario has a "call reason" field, calculate the indicators separately for at least the three or four main call types. Fixes are made by type too: wherever resolution is low, that is where the knowledge base is extended.

The limits of measurement

  • If resolution relies on the "outcome" field, the field itself can be filled wrongly
  • A repeat call may have another cause — a new question, another family member
  • If the sample check is small, the error rate depends on chance
  • Customer satisfaction is not directly in these indicators — a separate survey is needed
  • Indicators differ by call type — an overall figure hides the differences between types

Common mistakes

  • Tracking only the resolution rate
  • Trying to cut escalation without looking at reasons
  • No sample checks — the error rate is never measured
  • Merging all call types into one number
  • Substituting a vendor's "accuracy" figure for your own measurement

A legal and ethical note

These indicators measure the agent's work, not operators'. If AI indicators are used in evaluating operators, they can only be signals; the final decision must stay with a person. Using call recordings for review must follow the company's data policy and local requirements.

Data for measurement in Vexvon

In Vexvon AI Call Center every call's transcript, one-sentence summary and the fields you define in the scenario — for example "outcome" and "handoff reason" — stay in the panel. That lets you count resolution and escalation reasons without listening to calls one by one, and read transcripts quickly for sample checks. Calls, messages and website chats appear in one timeline per customer, so repeat contact can be tracked too; for overall reporting, see the analytics page.

How these indicators are used in a pilot is covered in the voice AI pilot plan, and their money value in voice AI ROI.

First step

Add two fields to your scenario: "outcome" and "handoff reason" — each with a closed list. A week later, calculate the five indicators for the first time. More articles are in the reliability & pilots section; to build the panel together, get in touch.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.