AI categorization validation: when a person should check
AI categories look convincing, but if a definition is vague, the number stays stable while counting the wrong thing. This guide covers AI categorization validation: two error types, a blind sample audit, when checking is mandatory, a calibration cycle and a checker's worksheet.
Short answer
A person must check the categories AI assigns to conversations in five situations: when a field is first switched on, when its definition or value list changes, when a value's share suddenly shifts, when a result will underpin an important decision, and for rare but serious values. Outside these, a small regular sample check is enough: 30–50 results per main field each week.
The aim of checking is not to "grade" the AI but to find two different errors: false positives (the conversation does not belong to the value, but AI put it there) and false negatives (the conversation belongs to the value, but AI put it elsewhere). Each has a different fix.
Why checking is needed
AI categories look convincing: a tidy label next to every conversation, precise percentages in the report. But the percentage is only as accurate as the labels. If a field's definition is vague, AI can choose wrongly in a consistent way, and the error never shows in the report: the number looks stable and sensible — it just counts the wrong thing.
That does not mean AI is "bad". Two people also read the same definition differently. Checking is there to show the definition works.
Two kinds of error
- False positiveAI put the conversation under "price objection", but the customer only asked the price and did not object. Result: the value's share is inflated.
- False negativeThe customer said "this is too expensive for me", but AI put it under "other" or another value. Result: the value's share shrinks and the problem stays hidden.
In plain terms, one is measured by precision (of the conversations AI said belong to the value, how many really do) and the other by recall (of the conversations that really belong to the value, how many AI found). Which matters more depends on the decision: for a rare, serious event (a legal complaint) missing none — recall — matters; for a trend report, a balance of both.
How to design a sample audit
- Selection10–15 random conversations from each value, plus 10–15 from "other". Sampling only the largest value hides errors in rare values.
- Blind readingThe checker first writes their own answer without looking at AI's choice, then compares.
- RecordingFor every mismatch: what AI chose, what the person chose, and why (vague definition, short conversation, missing context).
- CalculationFor each value, the share of correct choices and the conversations wrongly placed in "other".
When checking is mandatory
- First launch: every value is checked for the first 1–2 weeks.
- Definition or value changes: new results are checked after the change.
- A sudden share shift: a value's share jumps or drops sharply — check whether it is a real change or an analysis error.
- An important decision: a result that will underpin a budget, price, product or staffing decision.
- Rare, serious values: a person reads every case found.
A calibration cycle
- CheckA weekly sample audit.
- CauseSystematic errors are grouped: which two values get mixed, which phrasing AI misses.
- FixA "not included" line in the definition, a new example, or merging values.
- Re-analysisIf the definition changed significantly, the recent period's conversations are re-analysed.
- Re-checkThe same values are checked the next week — has the error dropped?
How much accuracy is enough
There is no universal threshold, and vendors' accuracy figures are usually measured on other data in other languages. Set your own threshold by decision: for trend tracking, what matters is that accuracy stays stable from period to period; if a budget decision rests on a value's share, that value deserves stricter checking; if individual customers will be contacted, every case found should be read.
Illustrative example
This is an illustrative example. A training centre switches on a "loss reason" field, and in the first week the "price" value accounts for half of losses. A blind audit shows several of 15 "price" results are really "schedule doesn't fit": the customer wrote "is there a weekend group at this price?" and AI chose on the word "price".
A boundary is added to the definition: "price — the customer objects to the price itself; questions about schedule, format or location do not belong here". After re-analysis the "price" share drops, the "schedule" value grows, and the decision changes: instead of a discount, a weekend group is opened.
The checker's worksheet
- Conversation link.
- The value AI chose and the evidence sentence.
- The checker's value (after blind reading).
- Agreement: yes / no.
- Cause of error: definition, short conversation, context, language, other.
- Proposed fix.
Typical mistakes
- Checking only in the first week, then forgetting.
- Checking only the largest values.
- Looking at AI's choice and saying "yes, that's right" — without blind reading, confirmation bias creeps in.
- Finding the error and not fixing the definition.
Limits
- A small sample only catches gross errors; measuring accuracy for rare values needs more conversations.
- Human checkers make mistakes too — two checkers and a clear definition reduce this.
- An AI category should not be the sole basis for a high-stakes decision about a customer or employee.
What Vexvon provides for checking
In Vexvon every result keeps the customer's own sentence and a link to that message, and the list of conversations behind each value can be opened — what a sample audit needs. How often phrases in "other" repeat is shown separately. After a definition or value list changes, old conversations can be force re-analysed; each analysis keeps a snapshot of the field definitions at that moment, so "why was this value chosen?" can be answered later. More: Vexvon analytics.
Next step
This week, blind-check 10 conversations per value in your most important field and record the mismatches. Write a boundary for the two values that get mixed most. For taxonomy rules, see conversation analytics taxonomy; we can set up the audit together in a demo.
Further reading on this topic: insufficient evidence analytics, analyze all customer conversations.
- Data, taxonomy & reliability6 min readConversation analytics CRM integration: linking call and chat data with the CRM
- Data, taxonomy & reliability6 min readConversation analytics taxonomy: how to build the right category system
- Data, taxonomy & reliability6 min readSpeech transcription accuracy in analytics: how errors distort results