System designCore3 min

Queueing & utilization

Why a server at 90% busy is five times slower than one at 50%, and why every healthy system keeps spare capacity on purpose.

1 · The idea

Every busy resource has a queue in front of it

A CPU, a thread pool, a disk, a database connection pool: each serves one piece of work at a time per slot. Work that arrives while all slots are busy waits. Your latency is the waiting plus the work.

Requests arriveburstyQueueServer10 ms eachλ per secondRequests arriveburstyQueueServer10 ms eachλ per second

Utilization ρ is arrival rate × work per request: 80 requests/s × 10 ms = 0.8, so the server is busy 80% of the time.

2 · The curve

Waiting explodes near full

For the simplest model (random arrivals, one server), time in the system is work ÷ (1 − ρ). At half busy you wait as long as the work itself; at 90% you wait nine times as long.

05010015020020406080Utilization (%)Response time (ms)the knee

From 50% to 80% busy, response time goes from 20 to 50 ms. From 80% to 95% it goes to 200 ms, for 15% more traffic.

The reason is burstiness: requests don't arrive evenly. A second the server spends idle is lost for good, but a burst has to wait its turn. The closer to full, the less idle time there is to soak bursts up.

3 · What to do about it

Keep headroom, and bound the queue

MoveWhy it works
Plan for 60–70% at peakStays left of the knee, with room for a spike or a lost server
Add servers (c slots)Several servers sharing one queue absorb bursts far better than one fast server
Bound the queueA full queue should reject fast (429/503) instead of making everyone wait
Time out waiting workA request the client already gave up on is pure waste to serve
Cut varianceSplit slow jobs off the fast path, so a 2 s export doesn't queue behind 10 ms reads

An unbounded queue is the trap: under overload it grows without limit, every request's latency grows with it, clients time out and retry, and the retries add more load. Bounded queues turn overload into fast, honest errors.

4 · In a real system

Find the knee under load

URL Shortener System Design

The peak load test walks up the curve

As traffic rises, the busiest component's utilization climbs and its latency bends upwards long before it errors. Watch which part reaches the knee first: that's where capacity goes next.

Check yourself

3 questions

1. A request needs 10 ms of work, and the server is 50% busy. Using work ÷ (1 − ρ), what's the response time?
2. Utilization rises from 80% to 90%. What happens to response time in that model?
3. Why reject requests when a queue is full, rather than letting it grow?

Takeaways

Remember this

  • Latency = waiting + work, and waiting grows like 1/(1 − utilization).
  • Past about 80% busy, a little more traffic costs a lot more latency.
  • Plan for 60–70% at peak; bound queues and reject early instead of queueing forever.