Skip to main content
Strategy
Blog

Chatbot Metrics: What to Measure and What to Ignore

Most chatbot dashboards put conversation count at the top. That number rises when you make the widget more insistent, so optimising for it is counterproductive. This article separates the figures that say something about the business from the ones that merely measure activity — and gives a formula for each. Support and sales are treated separately because their definitions of success are opposite, three figures that apply to every channel are set out, and the measurement traps — counting test conversations, using the mean instead of the median, not splitting by language — are named explicitly.

September 18, 20267 min read

Vanity numbers versus business numbers

A metric can be tested with one question: can this number be improved artificially, and does the business improve when you do? If the answer is 'yes, no', it is a vanity metric.

Conversation count is the classic example. Make the widget open on every page and the number doubles while sales stay flat. Response time can be a trap too: the bot answers in a second, so the average is always excellent and says nothing.

The useful figures share a property: the only way to improve them is to actually do the job better. The list below is built on that test.

Numbers to stop reporting

  • Total conversation count — measures traffic and how insistent the widget is, not value.
  • Average response time with bot and human counted together — the bot's replies hide the human queue, and a two-hour wait shows up as a ten-second average.
  • Messages sent — the worse a conversation goes, the larger this number becomes.
  • Containment rate without a quality check — a bot that frustrates people into leaving looks excellent here.
  • The bot's own confidence score — useful for debugging the model, with no place in a business report.
  • A sentiment score, if no decision is made on the basis of it.

Metrics for a support scenario

  1. Full resolution rateFormula: conversations closed without human involvement and not reopened by the customer within 24 hours, divided by all conversations. The reopening condition is essential — without it the number is measuring abandonment.
  2. Escalation rate, by reasonFormula: conversations reaching an operator divided by all conversations, broken down by trigger. The total says nothing; a rising knowledge-gap share means the material needs work.
  3. Time from escalation to first human replyFormula: the median between the operator notification and the operator's first message. Use the median rather than the mean — one long case distorts an average.
  4. Repeat contact rateFormula: cases where the same customer writes about the same subject a second time within seven days, divided by resolved conversations. This is the only number that shows whether 'resolved' meant resolved.
  5. Operator hours savedFormula: conversations resolved without a person multiplied by the average operator time per reply. Vexvon's reporting calculates this on a basis of roughly three minutes per reply — the methodology has to be stated, or the figure is marketing rather than evidence.

Metrics for a sales scenario

  1. Lead conversion rateFormula: conversations producing a routed lead divided by all conversations. There is no correct value — it depends on traffic mix — but a sharp change is always worth investigating.
  2. ContactabilityFormula: leads sales could actually reach divided by qualified leads. It exposes a flow collecting numbers nobody agreed to be called on.
  3. Sales acceptanceFormula: leads sales judged worth the call divided by leads contacted. The only real measure of qualification quality, and it requires sales to record a verdict.
  4. Time to first contactFormula: the median from the end of the conversation to the first outbound attempt. This number often matters more than qualification quality.
  5. Conversion by pathFormula: closed deals divided by leads, split between bot-originated and other. This is the figure that settles arguments about whether automation is costing you deals.

Three numbers that apply everywhere

  • Share of conversations arriving out of hours. Usually higher than teams expect, and the clearest justification for automation that exists.
  • Distribution by hour. The only figure that shows whether your rota matches actual demand.
  • Channel split. On its own it settles how much to invest in which channel.
  • AI cost per conversation. An AI bill you cannot break down is a monthly surprise; split per conversation it becomes a manageable line item.

Measurement traps

  • Counting test conversations. One testing week distorts the month entirely — they should be excluded by default.
  • Using the mean instead of the median. Conversation lengths and response times always have a long tail, and the mean hides it.
  • Counting bot and human replies together, which hides the human queue every time.
  • Not splitting by language. The language performing badly is usually the smallest by volume and is invisible in the total.
  • Reading only the successful conversations. Metrics tell you a problem exists; abandoned transcripts tell you where.
  • Turning a number into a team performance target, which converts it from something used into something manipulated.

What Vexvon measures

Reporting exports to Excel and covers an overview, working-hours analysis with an hourly heat map, channel split and lead analytics. The 'three shared numbers' section above comes directly from that report — hourly distribution and out-of-hours arrival are shown separately.

Conversations closed without a person and leads produced are measured separately, so support and sales metrics do not blur into each other. Operator hours saved is calculated with its methodology stated openly: roughly three minutes per reply. Test conversations are excluded from the figures by default.

AI cost and tokens are logged across seventeen distinct purposes with the model and channel recorded, and per call the tokens, latency, success rate and dollar cost are captured — which is what makes the 'AI cost per conversation' metric above possible rather than aspirational.

On the lead side, the record carries status, stage, priority, close reason and a contact-attempt count, with every change written to an activity log covering ten action types — that is where the verdict field required by sales acceptance and contactability lives. Conversations and question-answer pairs can be exported to Excel.

17Purposes AI cost is logged across
10Logged activity types
3Minutes per reply, saving methodology

Frequently asked questions

  1. Which single metric best shows chatbot performance?There is no single one. Full resolution rate with a quality check comes closest for support, and sales acceptance for sales. Conversation count replaces neither.
  2. Is containment rate a good metric?Only alongside a quality check. On its own it rewards a bot that frustrates people into leaving, because abandonment also looks like a conversation closed without a person.
  3. Why the median rather than the mean?Conversation lengths and response times have long tails. A handful of very long cases drags the mean, while the median describes what a typical customer experienced.
  4. Can support and sales metrics be combined?No. Their definitions of success are opposite: in support, a conversation not reaching a person is good; in sales, the right lead not reaching a person is a loss.
  5. How should AI cost be measured?Per conversation and broken down by purpose. An undifferentiated monthly figure cannot be managed, only be surprised by.
  6. How often should metrics be reviewed?The numbers monthly, the choice of metrics quarterly. Teams optimise for what is measured, so what gets measured is itself a decision.

Start with one metric

Do not begin by designing the whole dashboard. Pick one number — repeat contact rate for support, sales acceptance for sales — and measure it honestly for a month. That single figure almost always indicates what the next change should be, whereas a twelve-number dashboard usually indicates nothing.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demo

Your details are used only for the demo and to get in touch.

Book a Meeting with Vexvon

Pick a time that suits you in our calendar.