Skip to main content
Enterprise integration & API

AI integration sandbox: building a safe test environment

When you test on a live channel, the CRM fills with "Test Testov" leads and the automatic call rule rings a developer at night. This guide builds a sandbox for AI automation: five parts, test data, keeping tests out of reports, three steps from sandbox to live, checking AI changes, creating error cases and ownership.

October 7, 20266 min read

Short answer

A sandbox for AI automation is a separate environment where an integration is tried out without touching real customers or real systems. It has five parts: a test channel or test number, a test copy of the target systems (CRM, order system), a test knowledge base, separate test keys, and test data that is not real. After the sandbox comes a limited pilot — on a real channel but at small volume and under close watch — and only then full production. The core rules: test data must not reach live reports, test keys must not reach live systems, and the move to live happens against written criteria.

Why you cannot test in production

In teams without a sandbox, testing goes like this: a developer writes to the live WhatsApp number from their own phone, leads called "Test Testov" appear in the CRM, and the automatic call rule rings the developer at 2 a.m. In the worst case a test message reaches a real customer, or a test webhook changes a real order. At month end the report holds dozens of test leads, and nobody knows which ones are real.

A sandbox isolates these risks: you can make any mistake, because the consequences reach nobody.

The five parts of a sandbox

  1. Test channelA separate WhatsApp number, a test Instagram account, a test widget or a separate phone number.
  2. Test copies of target systemsTest environments of the CRM, order and payment systems — with no real data.
  3. Test knowledge baseA copy of the live base or a separate version, for checking new answers before they go live.
  4. Separate keysTest keys reach only test systems; live keys are not in the sandbox.
  5. Test dataInvented customers, invented numbers, cases prepared to match the scenarios.

Test data without real customers

The easiest route is copying the live database, but that creates an unnecessary copy of personal data. The better route is building test data from the scenarios: "new customer", "second enquiry from the same number", "enquiry with no number", "customer complaining", "question not in the knowledge base". Each gets an invented name and number. If real data is needed — for a volume test, say — it is anonymised: names, numbers and addresses are changed.

Keeping tests out of reports

When test data reaches live reports, the numbers go wrong: lead counts swell and response times look odd. The rule: every test record carries a "test" flag, and reports exclude flagged records automatically. In a sandbox, test data never reaches the live report at all; in a limited pilot the flag matters most, because test and real share the same system.

Three steps from sandbox to live

  1. SandboxEvery scenario is run with test data, and error cases are created artificially.
  2. Limited pilotA real channel but small volume: one number, one shift or a small share of customers. A sample is read every day.
  3. ProductionFull volume, normal monitoring and a rollback plan.

Each move to the next step happens against written criteria: "every scenario passed in the sandbox", "no serious error in a week of pilot". The business side of a pilot is covered in the omnichannel automation pilot.

Checking AI changes in the sandbox

A sandbox is not only for integrations. Changing the AI's instructions, adding a new document to the knowledge base or altering a routing rule should not be done live on a "let's see" basis either. A practical route: make the change in a test version first, check it against 20–30 prepared questions, compare old and new answers, and only then move it to live. Testing a change to a call flow is covered in testing a call flow.

Creating error cases on purpose

The sandbox's biggest advantage is being able to create errors safely. Switch off the test CRM and check the webhook retries. Break the API key and see the alert arrive. Send the same event twice and confirm no duplicate appears. These rules are covered in detail in API error handling and retry rules.

Owning and maintaining the sandbox

Sandboxes go stale quickly: the live system changes while the test copy stays the same, and tests start giving wrong results. So a sandbox needs an owner — usually IT — who regularly aligns the test environment with live, manages the test keys and knows who is testing when. Before major changes, the sandbox is refreshed.

An illustrative example

This is an illustrative example. A travel company is integrating an AI chatbot with its booking system. For the sandbox it prepares a separate WhatsApp test number, the booking system's test environment and 15 invented customers. Testing reveals that when the booking system does not respond, the bot wrongly says "no availability".

The bot is fixed to answer "I can't check right now; a manager will write to you" on an error. Then a limited pilot: one week with real customers for a single destination. With no serious errors, all destinations are opened.

Common mistakes

  • Testing on live channels and live systems.
  • Copying the full live database into the sandbox without anonymising it.
  • Test keys reaching the live system.
  • Not flagging test records — the report goes wrong.
  • Not refreshing the sandbox — tests pass against an old system.

Limitations

A sandbox does not fully reproduce real customer behaviour: real customers write unexpectedly, and real volume and networks differ. So a limited pilot after the sandbox is a must. Some external systems — social platforms, for example — have no full test environment; there, a separate test account and a careful limited pilot are used.

Testing in Vexvon

In Vexvon, before a channel goes live, the answers are checked with test conversations and the necessary fixes are made. Conversations held in the website widget's test mode are excluded from reports automatically. Call routing rules can be checked with a simulation before launch. The test environment for the company's own systems and the criteria for going live are agreed together in the integration plan. More on integrations.

Next step

For your next integration, fill in the five-part list: is there a test channel? a test copy of the target system? separate test keys? test data ready? are test records flagged? The "no" answers are the start of your sandbox plan. The testing section of the requirements document is shown in API integration requirements. More articles are in the enterprise integration section, and we can build your sandbox plan together during a demo.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.