Skip to main content
Support & knowledge base

Chatbot answer quality audit: scoring live conversations

A chatbot is tested before launch and then nobody looks at its answers — until a customer sends a screenshot of an old price. A post-launch audit is a different process: it checks, on real conversations, that the bot still works correctly. This article sets out a five-dimension rubric for a chatbot answer quality audit, random and risk sampling, a simple scoring scale, finding the source of a wrong answer, a backlog of fixes, the audit rhythm, a tighter check for high-risk topics and a one-week example.

September 28, 20267 min read

The bot is live. Are its answers any good?

A chatbot is tested before launch, goes live, and some time later the team stops looking at its answers. The report shows more conversations and fewer handovers — everything looks fine. Then one day a customer sends a screenshot of an old price, and it turns out the bot has been saying it for two weeks.

A pre-launch test shows whether the bot is ready. A post-launch audit shows whether it is still working correctly, and that is a different process. This article gives a framework for a chatbot answer quality audit: a scoring rubric, sampling, a scoring scale, turning findings into a backlog of fixes, and the audit's rhythm.

How it differs from pre-launch testing

  1. Pre-launch testRun on scenarios the team prepared. The question: does the bot answer the expected questions correctly?
  2. Post-launch auditRun on a sample of real customer conversations. The question: how does the bot handle real, unexpected questions, has the material gone stale, is the tone right, does escalation work?

The scenario families for pre-launch testing are in the chatbot test checklist. The audit's source is different — the messages customers actually write.

The QA rubric: five dimensions

Each answer is scored on five dimensions. They are kept separate because each is fixed in a different place.

  1. AccuracyIs the answer factually correct and in line with approved material? A wrong price, an old rule, an invented detail — the most serious error.
  2. CompletenessWas the customer's question fully answered? If one of two questions went unanswered, the answer is incomplete.
  3. EscalationWas a case that needed a person handed over? An unnecessary handover is an error too — the bot did not answer something it should know.
  4. ToneDoes the answer match the agreed brand tone? Coldness to an unhappy customer, a long reply to a simple question — tone errors.
  5. ResolutionDid the conversation end with an outcome for the customer — an answer, a next step, a handover? Or did the customer simply leave?

Sampling

Reading every conversation is impossible. The sample has to be chosen so that problems show up.

  • A random sample — 30–50 conversations a week, showing the overall level
  • A risk sample — negative sentiment, long conversations, the same question repeated
  • Handed-over conversations — they show why the bot stopped
  • Conversations from a new topic or campaign period — where material goes stale most
  • A share from each channel — the bot may behave differently on Instagram, WhatsApp and the website

A scoring scale

A simple scale makes audits comparable. Three levels per dimension are enough:

  1. 2 — goodThe dimension is fully met.
  2. 1 — acceptableA small shortfall with no harm to the customer: slightly long, one detail missing.
  3. 0 — errorA wrong fact, a missed escalation, an answer that confuses or annoys the customer.

Every answer scoring 0 on accuracy is flagged separately and fixed that week, whatever the average. The average is for the trend; a single wrong answer is for immediate action.

Finding the source of a wrong answer

When a wrong answer is found, the key question is: why? There are four typical sources:

  • Stale material — an old price or rule in the knowledge base
  • Contradictory material — two documents say different things
  • Retrieval found the wrong fragment — the material is right, but the bot took it from another topic
  • No rule — what the bot should do in this case was never written

In Vexvon, each answer shows which knowledge fragment it used. That cuts finding the source of a wrong answer to minutes: if the fragment is old — the material; if the fragment is correct but irrelevant — a retrieval problem.

The improvement backlog

An audit's output is a work list, not a report. Each finding is one line:

  • What was found — a short description of the conversation
  • Which dimension — accuracy, completeness, escalation, tone, resolution
  • Source — material, contradiction, retrieval, missing rule
  • Fix — a specific action
  • Owner and date
  • Check — did the same error recur in the next audit

Material-sourced fixes go to the knowledge base; how the material itself is checked is in the chatbot knowledge base audit.

The audit rhythm

  1. Weekly30–50 conversations, one person, one hour. Accuracy errors are fixed that week.
  2. MonthlyThe trend: the average on five dimensions, the most common source of error, the state of the backlog.
  3. Event-drivenA new campaign, a price change, a new product — a separate small audit in the first three days.

Example: a one-week audit

  1. Sample40 random and 10 risk conversations.
  2. FindingsTwo 0s on accuracy: an old delivery fee. One 0 on escalation: a customer writing «give me my money back» was not handed over. Five 1s on tone: answers unnecessarily long.
  3. FixesThe delivery fee is updated in the knowledge base; «refund» phrases are added to the escalation rule; the answer-length setting is moved to short.
  4. Next weekCheck whether the same errors recurred.

A tighter check for high-risk topics

On some topics one wrong answer costs more than dozens of good ones. A random sample is not enough for them.

  • The list: prices and discounts, legal terms, health, finance, safety
  • All conversations on these topics — or at least several every day — are read
  • The topic owner takes part: the head of sales checks prices, a lawyer checks terms
  • After every change — a new price, a new rule — the first 24 hours are checked separately

The auditor's checklist

  • Is every fact in the answer in the approved material?
  • Were all the customer's questions answered?
  • Was there a word or situation that required a handover, and was it handed over?
  • Does the tone suit the customer's mood?
  • Did the conversation end with an outcome?
  • If there is an error — which knowledge fragment was used?

What to measure

  • The average score on five dimensions — the weekly trend
  • The number of accuracy errors — the target is close to zero
  • Missed escalations
  • Fixes still open in the backlog, and their age
  • The same error recurring — the fix did not work

The Q&A export and sentiment share are in the analytics report — a good source for the risk sample.

Limits

An audit works on a sample and does not catch every error. Rare but costly errors — a wrong answer to a medical or legal question, for example — may never land in a random sample, so high-risk topics need a separate, tighter check. The auditor must know the domain — someone who does not know the product will not spot a wrong price.

Conclusion: quality is managed after launch too

A chatbot answer quality audit is not a one-off test but a weekly habit. A five-dimension rubric, random and risk samples, a simple scale, finding each error's source and a backlog — with these in place the bot does not degrade over time; it gets a little better every week.

Diagnosing repeat questions is covered in reducing repeat questions; the rest of the support section is in this category. To set up your audit process together, get in touch.

Frequently asked questions

  1. Who should audit chatbot answer quality?Someone who knows the product and the rules — usually the head of support or the knowledge owner. An outside auditor may not spot a wrong price or an old rule.
  2. How many conversations a week is enough?Usually 30–50 random and 10–20 risk conversations. At high volume, split the sample by channel.
  3. Can the audit be automated?Selecting risk conversations — yes. Scoring in full — no; it needs a person who knows what is correct.
Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.