API error handling and retry rules: how to set them up
When an error happens, a request either vanishes silently or the system retries it endlessly and knocks the other side over. This guide sets up API error handling: four error classes, retries with growing intervals, limits, a dead-letter queue, alert thresholds, a circuit breaker, errors mid-conversation and idempotency.
Short answer
API error handling starts with one question: does it make sense to retry this error? Temporary errors — server errors (5xx), timeouts, "too many requests" (429) — are retried with growing intervals. Permanent errors — a bad request, no permission, not found (most 4xx) — are not retried, because a retry will not change the result; they need to reach a person to be fixed. Every retry has a limit, a request that hits the limit is not lost but goes to a separate queue (dead-letter), and when a threshold is crossed the person responsible is notified. Retries are only safe together with idempotency.
Why you need an error rule
An integration without a rule behaves in one of two extremes. In one, a failed request simply vanishes: the lead never reaches the CRM and nobody knows. In the other, the system retries the error instantly and endlessly — if the other side is already overloaded, the retries knock it over completely, and the account may be blocked for exceeding the request limit.
An error rule sits between the two: what to retry, when, how many times, and when to stop and tell a person.
Error classes
- Temporary: retryServer errors (500, 502, 503, 504), timeouts, network drops.
- Rate limit: wait, then retry429 "too many requests" — wait as long as the other side says (the Retry-After header).
- Permanent: do not retry400 bad request, 401/403 permission, 404 not found, 422 validation failed.
- Unknown: carefullyAn unexpected response format — one or two retries, then a person.
Exponential backoff: growing intervals
Retries happen at growing intervals rather than equal ones: for example 30 seconds, 2 minutes, 10 minutes, 30 minutes, 2 hours. This is called exponential backoff. A small random addition (jitter) is added to the intervals so hundreds of requests that failed at the same moment do not all come back at the same moment. If the other side says how long to wait with a "Retry-After" header, that time is respected.
Limits: how many times and for how long
- Maximum attempts: for example 5–8.
- Maximum age: if an event is older than a set time (24 hours, say), it is no longer sent — stale data can do harm.
- Business meaning: a "lead created" notification an hour late is still useful; "customer on the line" is not.
- A request that reaches the limit goes to the dead-letter queue; it is not deleted.
The dead-letter queue
A dead-letter queue is where requests that have used up all their attempts are kept, with each request's data, the last error and the attempt history. Once the person responsible fixes the cause — renews a key, say — the requests are resent. If the queue never empties and nobody looks at it, it is just a tidier form of lost data: so the queue needs an owner.
Alert thresholds
Not every error is an alert — otherwise the team learns to ignore alerts. Practical thresholds: a permission error (401/403) immediately; the first request to land in the dead-letter queue; when the share of failed requests in the last hour crosses a set percentage; when the other system does not respond at all for a set time. Alerts go to a specific person or the on-call engineer, not a general group.
Circuit breaker: protecting the other side
If the other system is completely down, retrying every request loads it even as it recovers. The circuit breaker rule is: when consecutive failures cross a threshold, sending is paused temporarily, requests wait in a queue, after a while one test request is sent, and once it succeeds the queue is drained gradually.
An error mid-conversation: what the customer hears
When the AI calls an API during a conversation — for order status, say — the error happens in front of the customer. Retries here must be short: the customer will not wait 30 seconds. Write the rule in advance: one short retry, then an honest answer — "I can't check the status right now; an agent will write to you within 15 minutes" — and a notification to an agent. The AI must never guess a status when there is an error. The details of this design are covered in chatbot API and webhook integration.
Retries and idempotency
A timeout is the hardest case: the request may already have been processed on the other side, with only the response lost. A retry can then create a second record. So every operation that is retried must be idempotent — a second request with the same key must not create a second result. This is covered in detail in duplicate data between systems.
An illustrative example
This is an illustrative example. A training centre's integration sends leads to the CRM. One morning the CRM's API key expires and every request returns 401. The integration retries endlessly, the CRM account is blocked for an hour for "too many requests", and the leads are lost.
Under the new rule, 401 is not retried: IT is notified immediately and requests go to the dead-letter queue. After the key is renewed the queue is resent, and no lead is lost.
Common mistakes
- Retrying every error the same way.
- Instant retries with no interval.
- Endless retries with no limit.
- Nobody looking at the dead-letter queue.
- Retrying timeouts without idempotency.
Limitations
Even the best retry rule does not fix a long outage on the other side — it only ensures data is not lost and systems do not overload each other. The numbers (attempts, intervals, age) are chosen per business process and may differ for each integration. Write the rules into the requirements document — the section is shown in API integration requirements.
In Vexvon
Vexvon tracks the status of webhooks arriving from channels — pending, processing, retrying, failed and so on — and reprocesses webhooks that failed or got stuck. For webhooks Vexvon sends to the company's system, and for company APIs the AI calls, retries, limits and the answer given to the customer are agreed together in the integration plan. More on integrations.
Next step
Open the last month's error log for your integrations and sort the errors into the four classes. For each class, write down today's behaviour: is it retried, how many times, who is told? Webhook resends are covered in what is a webhook. More articles are in the enterprise integration section, and we can write your rules together during a demo.