System designCore3 min

Latency, throughput & percentiles

How long one request takes, how many you can serve at once, and why the slowest 1% is the number to watch.

1 · The idea

Two different kinds of fast

Latency is how long one request takes, from sending it to getting the answer. Throughput is how many requests the system finishes per second. A motorway makes the difference easy to see: the speed limit is latency, the number of lanes is throughput. Adding lanes doesn't get any single car home sooner.

ClientAPI4 ms of workDatabase6 ms20 ms there1 msClientAPI4 ms of workDatabase6 ms20 ms there1 ms

One request's latency is the sum along its path: network, queueing and work. Here about 20 + 4 + 1 + 6 + 1 + 20 = 52 ms.

You can often buy throughput with more machines. Latency is harder: it's set by the slowest step on the path, and by the distance the bytes travel.

2 · Measuring it

Averages hide the slow requests

Sort a minute's worth of response times. The p50 (median) is the one in the middle; the p99 is the one 99% of requests beat. The mean is dragged about by a few outliers and describes nobody.

p50
12 ms
mean
25 ms
p90
35 ms
p99
180 ms
p99.9
420 ms

A typical API: the median user waits 12 ms, but one request in a hundred waits 15 times longer.

The slow requests aren't rare people: a user who loads 40 resources per page hits a p99-slow one on a third of their page loads. That's why targets are written as percentiles: p99 under 200 ms, not average under 50 ms.

3 · The catch

Fan-out makes the tail everyone's problem

When one page waits for many backend calls, it's as slow as the slowest of them. If each backend is slow on 1% of calls, the chance that at least one of n calls is slow is 1 − 0.99ⁿ.

020406020406080100Backend calls per pagePages that hit a slow call (%)

With 100 backends, almost two page loads in three wait on somebody's p99.

Systems with wide fan-out fight this with hedged requests (send a second copy after the p95 and take whichever answers first), timeouts, and by trimming the tail of every dependency.

4 · Tying them together

Little's law

In a steady system, requests in flight = throughput × latency. It needs no assumptions about the traffic, so it's the quickest sanity check there is.

  1. Size a connection pool.

    2,000 requests/s that each hold a database connection for 25 ms need about 2,000 × 0.025 = 50 connections.

  2. Spot a slowdown.

    Latency doubles to 50 ms at the same traffic, and 100 connections are busy. A pool of 64 now queues, and latency climbs further.

  3. Find the ceiling.

    A server that can hold 200 requests at once, each taking 100 ms, tops out at 200 / 0.1 = 2,000 requests/s.

5 · In a real system

Watch the p99 move first

URL Shortener System Design

The redirect path's p99 under a peak

The URL shortener promises fast redirects at the p99. Run the peak load test and watch the latency climb as the API service fills up: the slow tail grows well before the service stops answering.

Check yourself

3 questions

1. You add a second, identical server behind the load balancer. What mostly changes?
2. An API has a mean of 25 ms, a p50 of 12 ms and a p99 of 180 ms. Which is true?
3. A service handles 500 requests/s and each takes 200 ms. How many are in flight at once?

Takeaways

Remember this

  • Latency is one request's time; throughput is requests per second. More machines mostly buy throughput.
  • Track percentiles (p50, p99), not the mean; the tail is what users feel.
  • Fan-out turns each backend's rare slow call into a common slow page.
  • Little's law: in flight = throughput × latency.