Chatbot answer quality audit: scoring live conversations
A chatbot is tested before launch and then nobody looks at its answers — until a customer sends a screenshot of an old price. A post-launch audit is a different process: it checks, on real conversations, that the bot still works correctly. This article sets out a five-dimension rubric for a chatbot answer quality audit, random and risk sampling, a simple scoring scale, finding the source of a wrong answer, a backlog of fixes, the audit rhythm, a tighter check for high-risk topics and a one-week example.
The bot is live. Are its answers any good?
A chatbot is tested before launch, goes live, and some time later the team stops looking at its answers. The report shows more conversations and fewer handovers — everything looks fine. Then one day a customer sends a screenshot of an old price, and it turns out the bot has been saying it for two weeks.
A pre-launch test shows whether the bot is ready. A post-launch audit shows whether it is still working correctly, and that is a different process. This article gives a framework for a chatbot answer quality audit: a scoring rubric, sampling, a scoring scale, turning findings into a backlog of fixes, and the audit's rhythm.
How it differs from pre-launch testing
- Pre-launch testRun on scenarios the team prepared. The question: does the bot answer the expected questions correctly?
- Post-launch auditRun on a sample of real customer conversations. The question: how does the bot handle real, unexpected questions, has the material gone stale, is the tone right, does escalation work?
The scenario families for pre-launch testing are in the chatbot test checklist. The audit's source is different — the messages customers actually write.
The QA rubric: five dimensions
Each answer is scored on five dimensions. They are kept separate because each is fixed in a different place.
- AccuracyIs the answer factually correct and in line with approved material? A wrong price, an old rule, an invented detail — the most serious error.
- CompletenessWas the customer's question fully answered? If one of two questions went unanswered, the answer is incomplete.
- EscalationWas a case that needed a person handed over? An unnecessary handover is an error too — the bot did not answer something it should know.
- ToneDoes the answer match the agreed brand tone? Coldness to an unhappy customer, a long reply to a simple question — tone errors.
- ResolutionDid the conversation end with an outcome for the customer — an answer, a next step, a handover? Or did the customer simply leave?
Sampling
Reading every conversation is impossible. The sample has to be chosen so that problems show up.
- A random sample — 30–50 conversations a week, showing the overall level
- A risk sample — negative sentiment, long conversations, the same question repeated
- Handed-over conversations — they show why the bot stopped
- Conversations from a new topic or campaign period — where material goes stale most
- A share from each channel — the bot may behave differently on Instagram, WhatsApp and the website
A scoring scale
A simple scale makes audits comparable. Three levels per dimension are enough:
- 2 — goodThe dimension is fully met.
- 1 — acceptableA small shortfall with no harm to the customer: slightly long, one detail missing.
- 0 — errorA wrong fact, a missed escalation, an answer that confuses or annoys the customer.
Every answer scoring 0 on accuracy is flagged separately and fixed that week, whatever the average. The average is for the trend; a single wrong answer is for immediate action.
Finding the source of a wrong answer
When a wrong answer is found, the key question is: why? There are four typical sources:
- Stale material — an old price or rule in the knowledge base
- Contradictory material — two documents say different things
- Retrieval found the wrong fragment — the material is right, but the bot took it from another topic
- No rule — what the bot should do in this case was never written
In Vexvon, each answer shows which knowledge fragment it used. That cuts finding the source of a wrong answer to minutes: if the fragment is old — the material; if the fragment is correct but irrelevant — a retrieval problem.
The improvement backlog
An audit's output is a work list, not a report. Each finding is one line:
- What was found — a short description of the conversation
- Which dimension — accuracy, completeness, escalation, tone, resolution
- Source — material, contradiction, retrieval, missing rule
- Fix — a specific action
- Owner and date
- Check — did the same error recur in the next audit
Material-sourced fixes go to the knowledge base; how the material itself is checked is in the chatbot knowledge base audit.
The audit rhythm
- Weekly30–50 conversations, one person, one hour. Accuracy errors are fixed that week.
- MonthlyThe trend: the average on five dimensions, the most common source of error, the state of the backlog.
- Event-drivenA new campaign, a price change, a new product — a separate small audit in the first three days.
Example: a one-week audit
- Sample40 random and 10 risk conversations.
- FindingsTwo 0s on accuracy: an old delivery fee. One 0 on escalation: a customer writing «give me my money back» was not handed over. Five 1s on tone: answers unnecessarily long.
- FixesThe delivery fee is updated in the knowledge base; «refund» phrases are added to the escalation rule; the answer-length setting is moved to short.
- Next weekCheck whether the same errors recurred.
A tighter check for high-risk topics
On some topics one wrong answer costs more than dozens of good ones. A random sample is not enough for them.
- The list: prices and discounts, legal terms, health, finance, safety
- All conversations on these topics — or at least several every day — are read
- The topic owner takes part: the head of sales checks prices, a lawyer checks terms
- After every change — a new price, a new rule — the first 24 hours are checked separately
The auditor's checklist
- Is every fact in the answer in the approved material?
- Were all the customer's questions answered?
- Was there a word or situation that required a handover, and was it handed over?
- Does the tone suit the customer's mood?
- Did the conversation end with an outcome?
- If there is an error — which knowledge fragment was used?
What to measure
- The average score on five dimensions — the weekly trend
- The number of accuracy errors — the target is close to zero
- Missed escalations
- Fixes still open in the backlog, and their age
- The same error recurring — the fix did not work
The Q&A export and sentiment share are in the analytics report — a good source for the risk sample.
Limits
An audit works on a sample and does not catch every error. Rare but costly errors — a wrong answer to a medical or legal question, for example — may never land in a random sample, so high-risk topics need a separate, tighter check. The auditor must know the domain — someone who does not know the product will not spot a wrong price.
Conclusion: quality is managed after launch too
A chatbot answer quality audit is not a one-off test but a weekly habit. A five-dimension rubric, random and risk samples, a simple scale, finding each error's source and a backlog — with these in place the bot does not degrade over time; it gets a little better every week.
Diagnosing repeat questions is covered in reducing repeat questions; the rest of the support section is in this category. To set up your audit process together, get in touch.
Frequently asked questions
- Who should audit chatbot answer quality?Someone who knows the product and the rules — usually the head of support or the knowledge owner. An outside auditor may not spot a wrong price or an old rule.
- How many conversations a week is enough?Usually 30–50 random and 10–20 risk conversations. At high volume, split the sample by channel.
- Can the audit be automated?Selecting risk conversations — yes. Scoring in full — no; it needs a person who knows what is correct.