System designCore3 min

Timeouts, retries, backoff & jitter

Retries turn brief blips into successes, and turn outages into disasters if done naively. Here's how to get the first without the second.

1 · Start here

Every remote call needs a timeout

Without a timeout, a call to a hung service waits forever, holding a thread, a connection and memory. Enough of those and the caller hangs too. Set the timeout from the dependency's p99 plus a margin, not from a round number like 30 seconds.

Better still, pass a deadline down the chain: if the user's request has 800 ms left, a call two hops down shouldn't start a 1-second wait.

2 · Retries

Retry what's transient, and only if it's safe

FailureRetry?Why
Timeout, connection resetYes, if the call is idempotentProbably a blip, but the first attempt may have worked
503, 429 (overloaded)Yes, after waiting (honour `Retry-After`)Hammering it now makes it worse
400, 404, validation errorsNoIt'll fail the same way every time
A non-idempotent write that timed outOnly with an idempotency keyOtherwise you may book the ticket twice

3 · The catch

Retries multiply down the stack

If every layer tries three times, a failing database at the bottom gets 3 × 3 × 3 = 27 attempts for each user click, exactly when it's struggling.

Client3 triesAPI3 tries eachService3 tries eachDatabase27 queries×3×9×27Client3 triesAPI3 tries eachService3 tries eachDatabase27 queries×3×9×27

Retry at one layer, usually the one closest to the user or the one that knows the call is safe, and cap retries with a budget: for example, retries may add at most 10% to a client's normal traffic. When the budget is spent, fail fast.

4 · Waiting between tries

Back off exponentially, and add jitter

Wait longer after each failure, doubling up to a cap: 100 ms, 200, 400, 800… That gives a struggling service room to recover. But if a thousand clients failed at the same moment, they'll all retry at the same moments too, in synchronized waves.

  • Backoff, no jitter
  • Full jitter
02004006008001,000050100150200Time after the failure (ms)Retries arriving (per 10 ms)

1,000 clients fail at once and retry after 100 ms. Without jitter they arrive as one spike; with a random wait between 0 and 100 ms they arrive spread out.

Exponential backoff with full jitter

typescript
const base = 100, cap = 10_000;
for (let attempt = 0; attempt < 4; attempt++) {
  try {
    return await call({ deadline });
  } catch (e) {
    if (!isRetryable(e) || retryBudget.empty()) throw e;
    const ceiling = Math.min(cap, base * 2 ** attempt);
    await sleep(Math.random() * ceiling); // full jitter: anywhere from 0 to the ceiling
  }
}
throw new Error('gave up');

5 · In a real system

A retry storm, live

URL Shortener System Design

Retries finish off Cassandra

When the URL shortener's cache dies at peak, failed redirects are retried and the database's load climbs past what the original traffic alone would cause. Run the scenario and watch the retries.

Check yourself

3 questions

1. Which failure should you not retry?
2. Client, API and service each retry up to 3 times. How many attempts can one click cause at the database?
3. What does jitter fix that backoff alone doesn't?

Takeaways

Remember this

  • Every remote call gets a timeout from the dependency's p99, and a deadline passed down.
  • Retry only transient failures of idempotent calls, at one layer, within a budget.
  • Back off exponentially up to a cap, with jitter, so retries don't arrive in waves.