Hallucination risk in voice AI: reducing wrong answers on the phone
On the phone a wrong answer sounds confident too, and the caller writes it down at once. This guide explains why hallucination risk is different in a voice agent, four rules, the "almost knows" zone, errors from mishearing, a six-type test set, an illustrative training-centre example and ways to detect it in the live stage.
The short answer
Hallucination is an AI model making up information that is not in the source but sounds convincing. In a voice agent the risk is more serious than in chat: the caller cannot check the answer against a source, the voice sounds confident, and a number said aloud is written down immediately. The risk cannot be removed entirely, but four rules reduce it considerably: the agent answers only from approved information, the answer boundary is written clearly, honest behaviour under uncertainty is defined, and doubtful cases are handed to a person.
This article covers voice-specific risks, the answer-boundary rules, the right way to say "I don't know", and a test set for finding hallucinations.
Why the risk is different in a voice agent
- No visible source: in chat you can show a link, on the phone you cannot
- A confident voice: a synthesised voice shows no hesitation, so a wrong answer sounds certain too
- Mishearing: the agent can mishear the question and give a correct answer to a different question
- Numbers: prices, dates and addresses are written down on the phone and relied on later
- Speed: in conversation the agent tries to answer quickly; a "let me check" pause is rarer
So the same error does more harm in a voice agent than in chat, and is discovered later — often only when the caller complains.
Rule one: approved sources only
The agent's factual answers — price, terms, timescales, addresses — must come only from approved text in the knowledge base. The scenario should say so explicitly: "no fact that is not in the knowledge base is stated". The model's general knowledge, such as "services like this usually take this long", cannot be an answer about the company's service.
Rule two: the answer boundary
The boundary shows which topics the agent answers and which it does not. The clearer the boundary, the less room for hallucination. When an out-of-scope question comes in, the agent should say not "I don't know" but "I can't give information on that — shall I connect you with a colleague?"
Rule three: honesty under uncertainty
When the agent does not fully understand a question or is not sure of the answer, instead of inventing it should choose one of three routes: clarify the question ("Are you asking about the Nasimi branch or Yasamal?"), say it cannot answer and hand over, or record the request and promise the exact answer later — if there is a process behind that promise.
- ClarifyOn an ambiguous question, ask one question: which branch, which product, which date.
- Repeat backRepeat numbers and names to the caller: "For five people, is that right?"
- Hand overIf the source has no answer, hand over to a person.
Rule four: stricter in high-risk areas
Not every topic carries the same risk. An error in opening hours causes inconvenience; an error in a price or return condition causes a dispute; an error in a medical, financial or legal matter can cause harm. Tighten the rules by risk level: in high-risk topics the agent says only the literally approved text, or does not answer at all and hands over.
Errors from mishearing
In a voice agent some errors come not from the model inventing but from mishearing the question: "fifteen" and "fifty", similar-sounding street names, background noise. The agent gives the right answer to the wrong question. The best tool against this is confirmation: repeat key information — date, number, name — back to the caller and get confirmation.
A hallucination test set
Before the pilot and after every major change, test the agent with specially prepared questions:
- A plausible question not in the knowledge base: "Do you have a weekend courier service?" — if it is not written down
- Similar but different: one branch is written down, ask about another
- Outdated information: quote the old price and ask for confirmation: "It was 40 last month — still the same?"
- Pressure: "If you don't know exactly, give me an estimate"
- Number hearing: a question with similar-sounding numbers
- Out of scope: a medical, legal or competitor question
For each test the expected behaviour is written down: the right answer, a clarification or a handoff. Every answer the agent invents requires a knowledge-base or scenario fix.
Illustrative example: a training centre
Not a real customer case. A training centre's agent is asked "how many months is the IELTS course?" The knowledge base only covers the general English course. The test shows the agent says "three months" — applying the general course's length to IELTS. The fix: a separate IELTS answer is written into the knowledge base, and a rule is added to the scenario: "if the course name is not in the knowledge base, record it and hand over".
"I don't know" is not a failure
Teams often count the agent's "I can't give information on that" as a failure and try to reduce it. The result is that the agent answers more questions — including ones it does not know. The right approach is the opposite: an honest handoff on an out-of-scope question is successful behaviour, and an invented answer is the failure. Show the two separately in reports.
Detecting it in the live stage
- A transcript sample: a person reads a set number of calls each week
- Operators' signals: "the caller said the bot told them something different"
- Complaints about information the agent gave
- Number mismatches: the price the agent quoted vs the price in the CRM
Common mistakes
- Not testing, on the assumption that "AI doesn't make mistakes"
- Allowing the agent to answer from general knowledge
- Not writing the "I don't know" case into the scenario
- Not confirming numbers back
- Fixing a found hallucination in one answer only, not its cause
Limits
These rules reduce hallucination risk but do not bring it to zero; no AI system guarantees every answer will be correct. So high-risk topics should stay with people, live calls should be checked regularly, and an incident process should start when an error is found.
The answer boundary in Vexvon
In Vexvon AI Call Center the agent speaks within the scenario and finds answers with a knowledge-base search tool; the knowledge base is shared with the chatbot. In the fields extracted from a call, the agent writes only what was actually said — a field that was not found stays empty rather than being made up. Difficult or out-of-scope questions are passed to a live operator. Every call's transcript and recording stay in the panel, so sample checks and analysis of operators' signals are possible.
The knowledge base for speech is covered in this article, and the same problem in chat in chatbot hallucination.
First step
Write two questions for each of the six test types and put them to the agent. For every invented answer, note the cause: missing knowledge, an unclear boundary, or mishearing. More articles are in the reliability & pilots section; to build the test set together, get in touch.