Skip to main content
Audio analysis

Speech analytics pilot plan for existing call recordings

The most common way to try AI call analysis is to upload a few calls, look at the output and say "looks sensible". That is not an evaluation: nothing says what it is compared with or what success means. This article builds a seven-step pilot plan on your existing recordings: choosing one goal, writing the call standard, collecting a representative sample, preparing a human reference set, writing success criteria in advance, running the pilot, and deciding whether to expand, limit or stop.

September 29, 20267 min read

Short answer

The most reliable way to start AI analysis on your existing call recordings is a small pilot with a written goal: one call type, a few dozen carefully chosen calls rather than several hundred, a reference set scored by people in advance, and success criteria agreed before the pilot begins. At the end, the answer should not be "we liked it" but "the criteria were met, not met or partly met".

The point of a pilot is not to prove how good a tool is. It is to learn, before rolling out to the whole team, what works and what does not on your calls, in your languages, against your criteria.

Step 1: pick one goal

"Analysing calls" is not a goal. A pilot should answer one specific question:

  • Cut the time a QA specialist spends per call — AI does the first pass, a person listens only to doubtful cases
  • Find a problem sampling misses — for example, mandatory information not being given
  • Collect evidence-backed examples for training new agents
  • See how a call standard is actually applied in real calls

The goal defines the success measure. If the goal is saving time, the measure is time. If it is finding problems, the measure is the share of flagged cases confirmed by a human review.

Step 2: write the call standard

AI scores a call against your criteria. If the criteria are unclear, so is the result. For a pilot, criteria should be simple and checkable:

  1. One question per criterion"Did the agent clarify the need and tailor the offer to it?" becomes two questions.
  2. Observable behaviourNot "was polite" but "asked for the customer's name and, at the end, asked whether they had any other questions".
  3. A not-applicable ruleFor each criterion, write when it does not apply: if there was no objection, the objection-handling criterion does not apply.
  4. No tone-of-voice criterionTranscript-based evaluation does not measure tone and intonation reliably. Google's Quality AI guidance also advises avoiding such questions.

Step 3: a representative sample of calls

The most common mistake is taking only "good" or only "problem" calls. The sample should reflect everyday reality. The split below is an illustrative example:

  1. Call typePick one call type — for example, inbound sales calls. Do not mix sales and support in a pilot.
  2. AgentsSeveral agents with different experience; not just the best or the weakest.
  3. LanguageThe languages your team uses, in proportions close to reality, including mixed-language calls.
  4. Recording qualityNot only clean recordings — include a few noisy, short or broken calls too. You need to know how they are handled.
  5. OutcomeCalls that ended in a sale and calls that did not; calls with and without a complaint.

There is no universal sample size. Google requires at least 100 examples per question and at least 40 per answer option to tune its own Quality AI model — that is a requirement of their product, not a norm for other tools. A practical approach for a pilot: make sure every important criterion has enough examples of both "yes" and "no".

Step 4: a human reference set

A pilot cannot be evaluated without knowing what the AI result is compared with. The reference set is the same calls scored by people against the same criteria.

  • Two people score each call independently, without seeing each other's answers
  • Criteria where the two people differ are noted separately; these are often badly written criteria
  • Differences are discussed and an agreed answer is recorded — that is the reference
  • A human answer is not automatically treated as flawless; that is why the agreement step exists

Step 5: write success criteria in advance

Success criteria must be written before the pilot starts, not after seeing the results. The ones below are an illustrative example; set the thresholds for your own goal:

  1. Agreement per criterionFor each criterion, the share of calls where the AI answer matches the reference. Not one overall average — criterion by criterion, because one may be very good and another weak.
  2. Error type on critical criteriaIf AI said "met" on a critical criterion and the human said "no", that is the most dangerous error and is counted separately.
  3. Share of unanalysable callsHow many calls could not be scored reliably because of audio quality, language or speaker separation.
  4. Specialist timeHow long reviewing one call takes with and without the AI result.

Step 6: run the pilot

  1. Prepare the recordingsExport the selected calls from telephony and give the files readable names — date, call type, a non-personal identifier.
  2. AnalyseUpload the calls, choose the call standard, wait for results. Note failed files with the reason.
  3. CompareFor each call, check the AI result against the reference criterion by criterion. Where they differ, read the transcript: is the cause the criterion's wording, a transcription error, or the AI's interpretation?
  4. Fix the criteriaRewrite badly worded criteria and analyse the same calls again. This loop usually runs a few times.

Step 7: the decision

  1. ExpandThe success criteria were met. Next: a second call type or more agents, with human review continuing as a sample.
  2. Limited useSome criteria are reliable and some are not. The reliable ones are used; the rest stay with people or are rewritten.
  3. StopAgreement is weak on the key criteria, or most calls cannot be analysed — for example, because of recording quality. Expanding without fixing the cause makes no sense.

Typical mistakes

  • Writing the success criteria after the pilot
  • A pilot without a reference: "the results look sensible" is not an evaluation
  • Choosing only clean, good recordings — real use brings surprises
  • Using pilot results straight away to evaluate agents
  • Starting too wide: five call types, twenty criteria, three languages at once

Limits

  • Agreement found in a pilot holds only for that call type, language and recording quality
  • Consent and privacy requirements for using call recordings depend on the country and must be checked before the pilot
  • A poor or incomplete recording must not count as a result against the agent
  • During the pilot, AI scores should not be used for decisions about agents

A pilot with Vexvon Audio Analyzer

  • Existing recordings are uploaded as files — mp3, wav, m4a, ogg, opus, flac, amr, wma and other formats; up to 10 files per upload, up to 200 MB each
  • Uploading the same file twice is detected and flagged
  • You write your call standard as text; on every call each step is judged as met, partial, missed or not applicable, with a comment and transcript lines
  • After you fix the criteria, the same calls are re-analysed from the existing transcript — without transcribing the audio again
  • A failed file is shown with the reason and can be re-queued without uploading it again

More: Vexvon Audio Analyzer. How recording quality affects the result is covered in the articles of the audio analysis implementation and reliability section.

First step

Write a one-page pilot plan: goal, call type, sample split, who builds the reference and the success criteria. Then choose 30–50 calls and give them to two people to score independently. To run a pilot on your own recordings, get in touch.

Further reading on this topic: call center quality assurance software, speech analytics accuracy testing, call recording quality speech analytics.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.