Duplicate data between systems: idempotency and IDs
One enquiry arrives once, but the integration creates it twice in the target system — and everyone assumes the customer wrote twice. This guide covers the causes of technical duplicates and five rules: external IDs, upsert, idempotency keys, preventing sync loops, a reconciliation report, an example and ownership.
Short answer
Duplicate data between systems is prevented with five technical rules: store each record's identifier from the source system (an external ID) in the target system; use an "update if it exists, otherwise create" (upsert) operation instead of plain create; give every event a unique idempotency key so a resend does not create a second record; in two-way sync, mark where a change came from so no endless loop forms; and run a regular reconciliation report. These rules do not solve business duplicates — one person writing from three channels; they guard against duplicates the systems themselves create.
Two kinds of duplicate
The first kind is a business duplicate: the same customer writes on Instagram and WhatsApp, and two records appear. Intake rules prevent that — they are covered in duplicate lead prevention. The second kind is technical: one enquiry arrived once, but the integration created it twice in the target system. This article is about the second kind.
Technical duplicates are more dangerous because they are invisible: the business assumes "the customer wrote twice", while the problem sits in the integration and repeats every day.
Where technical duplicates come from
- Resends: the sender got no reply and sent the same event again, and the receiver created both.
- Two-way sync loops: system A updates B, B sends the change back to A as "new", A sends it again.
- Parallel imports: the same file is uploaded twice, or by two people.
- No shared key: the target system cannot tell whether an incoming record arrived before.
- Two processes at once: two workers handle the same event at the same moment.
Rule 1: the external ID
When each record reaches the target system, its identifier in the source system is stored in its own field: "source: chat platform, ID: 48213". The next time the same ID arrives, the target system does not create a new record — it finds the existing one. Phone and email can also be keys, but they change; an external ID does not. The most reliable approach uses both: the external ID as the main key, the phone as a backup check.
Rule 2: upsert
Instead of a "create" operation, use "update if it exists, otherwise create". The target system first searches by key: if the record exists, it updates it; if not, it creates it. Many CRMs offer this out of the box — HubSpot, for example, updates a contact with the same email on import rather than creating a new one. Without upsert, every resend means a new record.
Rule 3: the idempotency key
Idempotent means doing the same operation once or five times gives the same result. To get there, each event gets a unique key — the event's ID, for example. The receiving system remembers the keys it processed recently, and when the same key arrives again it does not repeat the operation; it simply replies "already processed". Webhook resends are discussed in detail in what is a webhook.
Rule 4: preventing sync loops
In two-way sync, each change carries its origin: "this change came from system B". When system A receives a change from B, it does not send it back to B. On top of that, each field should have an "owner" — only one system may change it, the other reads it. Field ownership is covered in the CRM integration data mapping checklist.
Rule 5: a reconciliation report
Even the best rules do not catch every case, so regular checks are needed. Once a week, compare the two systems' counts: how many leads were created in the source this week, and how many in the target. If there is a gap — either way — the cause is investigated. Two records with the same external ID in the target are direct proof that one of the rules is not working.
An illustrative example
This is an illustrative example. An online shop passes leads from its chat platform to the CRM by webhook. The head of sales notices some customers appear twice in the CRM. Investigation shows the CRM sometimes replies late, the chat platform resends the event when it gets no reply, and the CRM creates both as new leads.
The fix: an "external ID" field is added in the CRM, the integration uses upsert instead of create, and the receiving side remembers processed event IDs for 24 hours. The weekly reconciliation report shows the gap dropping to zero.
Common mistakes
- Treating a technical duplicate as "the customer wrote twice".
- Matching on name alone.
- Setting up retries without idempotency.
- Syncing every field both ways.
- Deleting duplicates by hand without fixing the cause — they come back tomorrow.
Limitations
These rules depend on what both systems support: if the target system offers no upsert or external ID field, an intermediary service may be needed. Cleaning up existing duplicates is separate work — covered in CRM duplicate records. In a multi-system environment full consistency is rarely achievable, so the reconciliation report should be permanent.
Who owns the problem
Technical duplicates belong to whoever owns the integration — usually IT or the party that built it. When business users see a duplicate they should not delete it but report it to the integration owner: every duplicate is evidence for finding the cause. The reconciliation report also goes to the integration owner, and who investigates a gap is agreed in advance.
Duplicates in Vexvon
In Vexvon, if an open lead exists in the same company for the same number, the new lead is linked to it as a duplicate; CSV imports check for repeated numbers during upload; and automatic call rules have a company-wide dedupe window. The status of webhooks arriving from channels is tracked, and failed ones are reprocessed. For webhooks Vexvon sends to the company's system, upsert and idempotency rules on the receiving side are agreed together in the integration plan. More on integrations.
Next step
Run one query in your target system: how many records share a phone number or source ID, and how many seconds apart were they created? A gap of a few seconds is the sign of a technical duplicate. More articles are in the enterprise integration section, and we can check your integration together during a demo.