System designCore3 min
Timeouts, retries, backoff & jitter
Retries turn brief blips into successes, and turn outages into disasters if done naively. Here's how to get the first without the second.
1 · Start here
Every remote call needs a timeout
Without a timeout, a call to a hung service waits forever, holding a thread, a connection and memory. Enough of those and the caller hangs too. Set the timeout from the dependency's p99 plus a margin, not from a round number like 30 seconds.
Better still, pass a deadline down the chain: if the user's request has 800 ms left, a call two hops down shouldn't start a 1-second wait.
2 · Retries
Retry what's transient, and only if it's safe
| Failure | Retry? | Why |
|---|---|---|
| Timeout, connection reset | Yes, if the call is idempotent | Probably a blip, but the first attempt may have worked |
| 503, 429 (overloaded) | Yes, after waiting (honour `Retry-After`) | Hammering it now makes it worse |
| 400, 404, validation errors | No | It'll fail the same way every time |
| A non-idempotent write that timed out | Only with an idempotency key | Otherwise you may book the ticket twice |
3 · The catch
Retries multiply down the stack
If every layer tries three times, a failing database at the bottom gets 3 × 3 × 3 = 27 attempts for each user click, exactly when it's struggling.
Retry at one layer, usually the one closest to the user or the one that knows the call is safe, and cap retries with a budget: for example, retries may add at most 10% to a client's normal traffic. When the budget is spent, fail fast.
4 · Waiting between tries
Back off exponentially, and add jitter
Wait longer after each failure, doubling up to a cap: 100 ms, 200, 400, 800… That gives a struggling service room to recover. But if a thousand clients failed at the same moment, they'll all retry at the same moments too, in synchronized waves.
- Backoff, no jitter
- Full jitter
1,000 clients fail at once and retry after 100 ms. Without jitter they arrive as one spike; with a random wait between 0 and 100 ms they arrive spread out.
Exponential backoff with full jitter
const base = 100, cap = 10_000;
for (let attempt = 0; attempt < 4; attempt++) {
try {
return await call({ deadline });
} catch (e) {
if (!isRetryable(e) || retryBudget.empty()) throw e;
const ceiling = Math.min(cap, base * 2 ** attempt);
await sleep(Math.random() * ceiling); // full jitter: anywhere from 0 to the ceiling
}
}
throw new Error('gave up');5 · In a real system
A retry storm, live
URL Shortener System Design
Retries finish off Cassandra
When the URL shortener's cache dies at peak, failed redirects are retried and the database's load climbs past what the original traffic alone would cause. Run the scenario and watch the retries.
Check yourself
3 questions
Takeaways
Remember this
- Every remote call gets a timeout from the dependency's p99, and a deadline passed down.
- Retry only transient failures of idempotent calls, at one layer, within a budget.
- Back off exponentially up to a cap, with jitter, so retries don't arrive in waves.