Insufficient evidence in analytics: why you need an "unknown" status
A system that forces a value onto every conversation makes the report look fuller than reality. This guide covers insufficient evidence in analytics: the role of an "unknown" status, the three statuses, calculating percentages correctly, a coverage report and the honest answer to questions about uncollected data.
Short answer
Conversation analytics needs an "unknown" (or "unclear") status, because in many conversations there is simply no information for a field: the customer does not mention a budget, does not object, does not state an intent. If the system forces a value onto every conversation, the gap is filled with the nearest value, and the report looks fuller and more certain than reality. An "unknown" status prevents that false precision.
Three situations must be kept apart: "other" (there is information, but no matching value on the list), "unknown" (there is no information) and "not analysed" (the conversation has not been analysed yet, or failed). When the three are lumped together, neither the list's quality nor the report's coverage is visible.
The cost of forced categorisation
Imagine a "budget range" field with only four values and no "unknown". Most customers say nothing about budget. AI is forced to pick a value for every conversation and usually picks the most "middle" one. The report then shows most customers have a mid-range budget — when it is really just missing information piled into one value. Marketing builds mid-segment ads on that number.
Three statuses: definition and difference
- OtherThe customer said something about this field, but it fits none of the listed values. Message: the list may need extending.
- UnknownThe conversation holds no information for this field. Message: the topic never came up.
- Not analysedThe conversation is still queued, does not meet the analysis conditions, or failed. Message: the number is incomplete.
"Unknown" is an insight in itself
A high "unknown" share is not a problem but information. If budget is unknown for 70% of leads, the sales team is not asking about budget — and perhaps should. If "loss reason" is unknown in most conversations, customers leave without giving a reason — a follow-up question is needed. Tracking the "unknown" share per field is a signal for both conversation quality and the script.
Reducing the "unknown" share the right way
If the "unknown" share is high, it must not be reduced by telling AI to "choose more boldly" or by deleting the "unknown" value — that is just a return to forced categorisation. The right way is to create the information in the conversation itself: add one or two clarifying questions to the agent's or bot's first messages ("which area are you looking in?", "when do you need it?"). When customers answer, the information appears in the conversation and the field fills honestly. A falling "unknown" share after that change shows both the analytics and the sales conversation have improved.
How to calculate percentages in reports
There are two different percentages, and mixing them creates a false impression. The first is relative to all conversations: "price objections in 12% of conversations with an objection". The second is relative to conversations that have the information: "40% of customers whose budget is known are in the lower segment". Write the denominator next to every percentage and show the "unknown" share separately. Otherwise a figure read as "40% of customers" actually applies to a small fraction of them.
A coverage report
- How many conversations there were in the period.
- How many were analysed, how many are queued, how many failed or were skipped.
- Among those analysed, the "unknown" share for each field.
- The "other" share.
These four lines belong at the top of every report. They show readers immediately how complete the numbers are.
Questions about data that was never collected
The most extreme form of "unknown" is a field that was never set up. A manager asks "which colour do customers ask for most?", but no "colour" field exists. A poor system takes a nearby field — say, "product of interest" — and invents an answer. A good system says plainly: this information is not collected; a field needs to be set up. That is not a pleasant answer for the user, but it is the only honest one.
Letting AI choose "unknown"
State clearly in every field's instruction: "if the conversation has no information on this, choose unknown; do not guess". Also give examples of which conversation should be "unknown" and which "other". During checks, record cases where "unknown" was chosen wrongly (the information was there but AI missed it) — part of AI categorization validation.
Illustrative example
This is an illustrative example. A real estate agency switched on a "number of rooms" field without "unknown". The report shows most customers are looking for two-room flats, and the agency focuses on two-room projects. When "unknown" is added and conversations are re-analysed, it turns out most customers never mention the number of rooms — AI had been putting them into the most common value.
The decision changes: first, "how many rooms are you looking for?" is added to agents' first message, and the real split is measured again a month later.
Typical mistakes
- Leaving "unknown" out of a field.
- Merging "unknown" and "other" into one value.
- Showing percentages over a denominator with "unknown" removed without saying so.
- Counting unanalysed conversations as zero results.
- Hiding the "unknown" share as a problem — it is a signal in itself.
Limits
- AI sometimes chooses "unknown" when the information is there — checks catch it as a false negative.
- If "unknown" is very high, the field may not suit that channel or conversation type.
- The coverage figure shows how complete the analysis is, not that the result is correct.
"Unknown" and coverage in Vexvon
In Vexvon every field returns a list, and when no information is found in the conversation the list stays empty — AI is not asked to guess. Answers that fall into "other" are kept separately. Reports from the analysis base show how much of the period has been analysed, and the panel counts analysed, queued and failed conversations. When a manager asks about data the company does not collect, the AI assistant does not substitute a nearby field; it says the data is unavailable. More: Vexvon analytics.
Next step
Check that every field has separate "unknown" and "other" values, and add a four-line coverage block at the top of your report. For taxonomy rules, see conversation analytics taxonomy; we can build the report format together in a demo.
Further reading on this topic: conversation analytics dashboard metrics.
- Data, taxonomy & reliability6 min readConversation analytics CRM integration: linking call and chat data with the CRM
- Data, taxonomy & reliability6 min readAI categorization validation: when a person should check
- Data, taxonomy & reliability6 min readConversation analytics taxonomy: how to build the right category system