Conversation analytics taxonomy: how to build the right category system
Without a category system, month-to-month comparison breaks and the numbers inspire no trust. This guide sets out six rules for a conversation analytics taxonomy, the steps to build it, the right number of values, an illustrative example and typical mistakes.
Short answer
A good category system (taxonomy) for conversation analytics rests on six rules: each field answers one question; values do not overlap (they are mutually exclusive); every value has a definition and examples; "other" and "unclear" are separate values; a value's technical key never changes after it is created; and every change is recorded as a version. Without these rules, month-to-month comparison breaks and the numbers inspire no trust.
A taxonomy is not a document you build once and forget. It changes as customers change, but in a controlled way: a new value is added, an old one is switched off rather than deleted, and past results either stay as they are or are deliberately re-analysed.
Field, value, level
A field is one question: "what did the customer object to?". A value is one possible answer: "price", "timing", "competitor". A level is a hierarchy of values: "delivery" at the top, "delay" and "damaged parcel" below. Two levels are enough for many businesses; more than three usually makes both AI's and people's choices unstable.
Rule 1: one field, one question
The most common mistake is piling everything into a single field called "topic": product names, objections, complaints, intents. In such a field "price" can be a question topic, an objection and a complaint all at once, and the number means nothing. Each field should ask one question: "what did they ask about?", "what did they object to?", "why didn't they buy?". When fields are separate, cross-tabs (such as product × objection) become meaningful.
Rule 2: mutually exclusive values
A conversation that falls into one value must not also fall into another value of the same field — if the field is single-value. If "price" and "discount" are separate values, where does "is there a discount?" go? Such overlaps should either be merged or given a clear boundary in the definition. If a conversation can genuinely have several answers (two objections, for example), make the field multi-value — but the values still must not duplicate each other.
Rule 3: definition and examples
- NameShort, close to the customer's language: "Delivery time".
- DefinitionOne sentence: what belongs to this value.
- Not includedThe boundary with the nearest neighbouring value: "delivery cost does not belong here".
- Examples2–3 phrases customers really wrote, with misspellings.
Rule 4: "other" and "unclear"
Every field needs two special values. "Other" — the customer said something, but it fits none of the listed values. "Unclear" — the conversation holds no information for this field. They differ: the first shows the list is incomplete, the second that the information is missing. The share of "other" is the list's quality indicator: if it is above 15–20%, look at the repeating phrases there and add a new value.
Rule 5: a stable key, a changeable name
Every value has two names: the label people see ("Delivery time") and the technical key the system stores ("delivery_time"). The label can change at any time — a translation, clearer wording. The key must not change after creation, because past results are tied to it. If the key changes, this month's "delivery_time" and last month's "delivery_duration" are counted as different things and the trend breaks.
Rule 6: versioning and switching off
When a value is no longer needed, switch it off rather than delete it: it is no longer chosen in new conversations but still appears with its label in past results. If it is deleted, old reports are left with a gap or an unreadable technical key. Record every change in a log: date, what changed, why. When the value list changes significantly, consider re-analysing old conversations so old and new results do not mix.
Building the taxonomy: steps
- Decision questionsFor which decisions? Each decision creates one or two fields.
- ReadingRead 150–200 conversations by hand and write down candidate values.
- Draft6–15 values per field plus "other" and "unclear", with definitions and boundaries.
- TestTwo people categorise 50 conversations separately; where they disagree is where the definitions are weak.
- Launch and calibrateAnalysis with AI, sample checks in the first weeks, sharpening definitions.
How many values is right
Too few values (3–4) make results generic: "service problem" says nothing. Too many (30+) cause confusion between neighbouring values and leave few conversations per value. In practice, 6–15 values per field work well. If you need more, split into two levels: first a top category, then subcategories within it.
Illustrative example
This is an illustrative example. An insurance agency starts with 40 values in a "topic" field. Two weeks later, checks show AI and agents often choose differently between "policy price", "price calculation" and "discounts". The taxonomy is split into two fields: "insurance type" (motor, property, travel, health, other) and "question type" (price, terms, documents, payment, claim notification, other). Each field keeps 6–7 values.
Result: the "insurance type × question type" cross-tab gives a clearer picture than the previous 40 values, and agreement in checks goes up.
Typical mistakes
- Piling different questions into one field.
- Building categories around department names rather than customer language.
- Deleting values or changing their keys.
- Leaving out an "other" value — AI is forced to put everything into the nearest value.
- Rewriting the list every month.
Taxonomy in Vexvon
In Vexvon the company builds fields and values in its own panel: name, instruction, closed value list, single or multiple values. AI can only pick from the list; a value's key is created once from its label and does not change afterwards. A field or value with results is switched off rather than deleted; full deletion needs separate confirmation. Frequently repeated phrases that fall into "other" appear in a separate list, and after a value list changes, old conversations can be force re-analysed. More: Vexvon analytics.
Next step
Take your existing fields and check each one: does it ask one question, does it have "other", does every value have a "not included" line? To find topics from scratch, see customer inquiry analysis; we can build the taxonomy on your own conversations in a demo.
Further reading on this topic: AI categorization validation, insufficient evidence analytics.
- Data, taxonomy & reliability6 min readConversation analytics CRM integration: linking call and chat data with the CRM
- Data, taxonomy & reliability6 min readAI categorization validation: when a person should check
- Data, taxonomy & reliability6 min readSpeech transcription accuracy in analytics: how errors distort results