System designCore3 min
Queueing & utilization
Why a server at 90% busy is five times slower than one at 50%, and why every healthy system keeps spare capacity on purpose.
1 · The idea
Every busy resource has a queue in front of it
A CPU, a thread pool, a disk, a database connection pool: each serves one piece of work at a time per slot. Work that arrives while all slots are busy waits. Your latency is the waiting plus the work.
Utilization ρ is arrival rate × work per request: 80 requests/s × 10 ms = 0.8, so the server is busy 80% of the time.
2 · The curve
Waiting explodes near full
For the simplest model (random arrivals, one server), time in the system is work ÷ (1 − ρ). At half busy you wait as long as the work itself; at 90% you wait nine times as long.
From 50% to 80% busy, response time goes from 20 to 50 ms. From 80% to 95% it goes to 200 ms, for 15% more traffic.
The reason is burstiness: requests don't arrive evenly. A second the server spends idle is lost for good, but a burst has to wait its turn. The closer to full, the less idle time there is to soak bursts up.
3 · What to do about it
Keep headroom, and bound the queue
| Move | Why it works |
|---|---|
| Plan for 60–70% at peak | Stays left of the knee, with room for a spike or a lost server |
| Add servers (c slots) | Several servers sharing one queue absorb bursts far better than one fast server |
| Bound the queue | A full queue should reject fast (429/503) instead of making everyone wait |
| Time out waiting work | A request the client already gave up on is pure waste to serve |
| Cut variance | Split slow jobs off the fast path, so a 2 s export doesn't queue behind 10 ms reads |
An unbounded queue is the trap: under overload it grows without limit, every request's latency grows with it, clients time out and retry, and the retries add more load. Bounded queues turn overload into fast, honest errors.
4 · In a real system
Find the knee under load
URL Shortener System Design
The peak load test walks up the curve
As traffic rises, the busiest component's utilization climbs and its latency bends upwards long before it errors. Watch which part reaches the knee first: that's where capacity goes next.
Check yourself
3 questions
Takeaways
Remember this
- Latency = waiting + work, and waiting grows like 1/(1 − utilization).
- Past about 80% busy, a little more traffic costs a lot more latency.
- Plan for 60–70% at peak; bound queues and reject early instead of queueing forever.