Skip to main content
Reliability & pilots

Planning an AI voice agent pilot: a narrow start, a baseline and decision criteria

A pilot that tests everything at once cannot show why it failed. This guide covers the pilot's scope, a six-week plan — preparation and baseline, internal test calls, live running, analysis — a five-indicator scorecard, decision criteria, an illustrative delivery service example and common mistakes.

September 30, 20266 min read

The short answer

An AI voice agent pilot should start narrow, be measurable and end with a decision criterion written in advance. A practical frame is six weeks: one week of preparation and baseline, one week of internal test calls, three weeks live on part of the real calls, and one week of analysis and decision. Throughout the pilot the human fallback — the route to a person for calls the agent cannot handle — must always be open.

The pilot's aim is not to prove that "AI works" but to answer a specific question: for this call type, on this line, does the agent meet the defined criteria? This article sets out the plan week by week, a pilot scorecard and how the go/stop decision is made.

Pilot scope: narrow and clear

  • One or two call types — for example, opening hours and booking requests
  • One line or one number — not all of them
  • Specific hours — for example, out of hours or peak hours
  • One language — others in a later stage
  • A clear boundary: what the agent does not do and when it hands over

The narrower the scope, the clearer the result. A pilot that tests everything at once cannot show why it failed.

Week 1: preparation and baseline

  1. BaselineThe last four weeks' figures for the selected call type: volume, answer rate, waiting time, outcome.
  2. ScenarioWhat the agent will say, what it will ask and which fields it will record.
  3. KnowledgeAn approved text and an owner for every topic the agent will answer.
  4. FallbackUnclear question, request for a human, technical problem — where each one goes.
  5. RolesThe pilot owner, who reviews calls, who makes fixes.

Week 2: internal test calls

Before connecting real customers, the team makes its own test calls. At least 30–50: typical questions, two questions in one sentence, names and numbers that are easy to mishear, silence, a request for a human, an out-of-scope topic, background noise. Each test call goes into a table: expected behaviour, actual behaviour, fix.

Weeks 3–5: live

The agent takes real calls on the selected line and hours. In the first days look at every call's transcript, then move to a daily sample. Once a week run a fix cycle: problems found, knowledge-base and scenario changes, a check with a test call. Do not make many changes in one week — it becomes hard to tell which one worked.

In the live stage, operators' feedback is the most valuable source: what the caller said on a handed-over call, what the agent misunderstood. Set up a simple channel for it.

The pilot scorecard

Write down before the pilot: which indicator, which target, how it is measured. An illustrative example:

  1. Full resolution rateCalls the agent finished without handing over — the target is set against the baseline.
  2. Correct handoffHow many out-of-scope calls were handed over correctly — via sample checks.
  3. Observed errorsCalls where wrong information was given — target near zero; each one is analysed.
  4. Hang-upsCallers ending the call while talking to the agent — should be no worse than the baseline.
  5. OutcomeBookings, enquiries or orders — compared with the baseline.

Week 6: analysis and decision

The decision is one of three: expand (another call type, line or hours), continue and improve within the same scope, or stop. Make the decision on the scorecard and call samples, not on one or two memorable calls. Stopping is a result too: why it did not work is valuable information for the next attempt.

Decision criteria: an example

  • Expand: all targets met, errors analysed and fixed
  • Continue: some targets met, the causes of problems are clear
  • Stop: observed errors keep recurring, or the fallback does not work
  • Stop: the team has no time to support the pilot — a resource problem, not a technology one

Illustrative example: a delivery service

Not a real customer case. A delivery service ran its pilot on out-of-hours calls and only for status and opening hours questions. Baseline: every night-time call went unanswered. In the live stage the agent took night calls, answered status questions with the tracking link and recorded other requests for the morning.

In the fourth week a review showed that the agent treated "the courier didn't come" complaints as status questions. A separate rule was added to the scenario: calls in a complaint tone go to the morning duty officer with priority. Decision in week six: expand to daytime peak hours.

Who to tell about the pilot, and how

Operators need to know in advance when the pilot starts, which calls move to the agent and where to log signals. Management should get a short weekly report: a one-table scorecard, three call samples and next week's changes. A separate announcement to customers is usually unnecessary, but it matters not to hide that the agent is AI, and to keep the recording notice.

After an "expand" decision

Expansion should be planned as a new pilot: a baseline, test calls and a scorecard for the new call type or line. A scenario that worked in the first pilot may not give the same result on a new call type. The safest route is one variable at a time: new hours, or a new call type, or a new line — but not all at once.

Common mistakes

  • Starting without a baseline
  • Opening the pilot on every line at once
  • Writing the decision criteria after the pilot
  • Not testing the fallback
  • Not appointing a pilot owner

Limits

Six weeks is not the same for every business: with low call volume, a longer period may be needed for a statistically meaningful result. Seasons and campaigns can distort the result — note them. If the pilot runs with real customers, the recording notice and data processing rules must be checked in advance.

A pilot with Vexvon

Vexvon AI Call Center works on top of your existing SIP or PBX line, so there is no need to change numbers for the pilot. Routing rules keep the pilot narrow: only a specific number, or only out-of-hours calls, reach the agent, while the rest go as before; you can check the rule in the simulator. Every call's transcript, recording, summary and extracted fields stay in the panel — that is where the samples for the scorecard come from. A difficult call is passed to a live operator.

The call type for the pilot is covered in the selection matrix, and the economics in voice AI ROI.

First step

Write a one-page pilot plan: call type, line, hours, baseline figures, five indicators and decision criteria. That page is the pilot's contract. More articles are in the reliability & pilots section; to plan the pilot together, get in touch.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.