How to test a call flow before you change it
A new night rule is added — and VIP customers suddenly land in the night scenario. This guide covers why a "small change" is dangerous, six test groups, a five-column test table, testing routing rules, a regression set, audio quality tests, the test process and an illustrative example.
The short answer
Every change to a call flow — a new menu option, business hours, scenario wording, a knowledge-base answer — can break something elsewhere. So six groups of scenarios should be tested before a change: a normal call, wrong input, silence, a request for a human, business hours and exceptions, and overload. For each group the expected result is written in advance and compared with a test call.
The most important part of testing is not the change itself but the parts it does not touch — "regression testing". When you add a new holiday rule, you need to check the VIP rule still works. This article sets out the test table, the regression set and who owns the test process.
Why a "small change" is dangerous
A call flow is made of interconnected rules. If rules are checked by priority, a new rule added near the top can "hide" several rules below it. Changing one sentence in the greeting can affect the agent's later behaviour. Updating one knowledge-base answer can change the answer to a similar question. So "we only changed one line" is no reason to skip testing.
Six test groups
- 1. NormalA typical caller's typical question: hours, price, booking. Does the flow finish as expected?
- 2. Wrong inputWrong key, an unclear answer, changing one's mind mid-question, two questions in one sentence.
- 3. SilenceThe caller says nothing or pauses for a long time. How many times does the agent ask, and what then?
- 4. Request for a human"Operator", "a person", "0" — at every step. During and outside business hours.
- 5. Hours and exceptionsStart and end of business hours, weekends, holidays, time zone.
- 6. OverloadSeveral calls at once, operators busy — does the fallback route work?
The test table
Each test row has five columns: scenario, input (what is said or pressed), time and number, expected result, actual result. The expected result is written before the test — written afterwards, it is easy to say "that's fine too".
- Scenario: public holiday, a call within normal business hours
- Input: "Are you open tomorrow?"
- Time and number: 1 January, 11:00, sales line
- Expected: the agent states the holiday schedule and does not hand over to an operator
- Actual: recorded during the test
Testing routing rules
Routing rules are the part most often got wrong, because they are combinations: time, dialled number, caller number. Write at least three tests per rule: a call that matches, one that does not, and one on the boundary (for example, one minute before business hours end). If there is a simulator, checking "where would this call land?" for every combination without making real calls is the fastest route.
The regression set
A regression set is a fixed set of tests repeated after every change. It can be 15–25 tests covering the flow's most important paths: the three most frequent call types, a request for a human, out of hours, the VIP rule, an emergency phrase. Every past incident is also added to the set as a test — so the same error does not come back.
Silence and audio quality
Make test calls in real conditions, not from a quiet office room: from the street, a car, a weak mobile signal. Audio quality affects what the agent hears, and a test that passes in the office may fail with a real caller. The silence test matters especially: the caller's phone is in their pocket and the call stays open — when and how does the agent end it?
The test process: who and when
- Before the changeThe person making the change writes the expected result.
- After the changeSomeone else runs the tests — the person who made the change may not see their own mistake.
- ResultIf all tests pass, the change goes live; if not, it is rolled back or fixed.
- RecordThe change, the test result and the go-live date stay in a log.
Illustrative example: a new night rule
Not a real customer case. A company adds a new rule for night hours: after 20:00 calls go to the night scenario. The rule is placed at the top of the list. The regression test shows that VIP customers now also land in the night scenario — previously they always went to the duty manager. The cause: the new rule sits above the VIP rule.
The fix: the VIP rule is moved back to the top, the test is repeated, and both rules work as expected. The case is added to the regression set as the test "20:30, VIP number → duty manager".
A rollback plan
Before every change, ask: if this does not work, how — and how quickly — do we get back to the previous state? The previous version of the rule, scenario or knowledge-base answer must be kept. Release changes at a low-volume hour — early in the morning, for example — and watch live calls for the first hour. If a problem shows up, roll back first rather than trying to fix it live, then analyse the cause calmly.
Common mistakes
- Skipping the test because it is "a small change"
- Writing the expected result after the test
- Testing only the changed part and not the rest
- Testing only from the office, with good audio
- Forgetting holiday and time-zone tests
Limits
A test set cannot cover every real case; callers will always say something unexpected. So testing must work alongside regular review of live calls. A conversational agent's answers to the same question may not be word for word identical each time — in tests, look at behaviour (correct information, correct handoff), not the exact words.
Testing a rule in Vexvon
In Vexvon AI Call Center routing rules are read from top to bottom by priority, and a rule can be tried in the simulator before it goes live: "where would a call from this number at this time land?" That lets you run the routing tests above without making real calls. Scenario and knowledge-base changes need test calls; every call's transcript and recording stay in the panel, so comparing the actual result with the expected one is easy.
A dedicated hallucination test set is covered in voice AI hallucination, and the process when an error is found in the incident process.
First step
This week, write a 15-test regression set: two or three tests from each of the six groups. Run it for the first time before your next change. More articles are in the reliability & pilots section; to prepare the test set together, get in touch.