System designCore3 min

Availability, SLAs & nines

What 99.99% available really allows, why a chain of services is less available than any link in it, and how redundancy wins it back.

1 · Counting nines

Each nine is ten times less downtime

Availability is the share of time, or of requests, that the system serves correctly. It's quoted in nines, and each extra nine cuts the allowed downtime by ten.

AvailabilityDown per yearDown per 30 days
99% (two nines)3.65 days7.2 hours
99.9% (three nines)8.8 hours43 minutes
99.95%4.4 hours22 minutes
99.99% (four nines)53 minutes4.3 minutes
99.999% (five nines)5.3 minutes26 seconds

Four nines leaves about four minutes a month: not enough time for a person to be paged, log in and fix anything. Past three nines, recovery has to be automatic.

2 · The vocabulary

SLI, SLO, SLA and the error budget

TermWhat it isExample
SLI (indicator)What you measureShare of requests that succeed in under 300 ms
SLO (objective)The target you hold yourself to99.9% of requests over any 30 days
SLA (agreement)A promise to customers, with a penalty99.5%, or a service credit. Set looser than the SLO
Error budgetWhat the SLO lets you miss0.1% of requests: spend it on releases and experiments

The error budget turns reliability into a decision: while there's budget left, ship; when it's spent, slow down and fix things.

3 · The catch

Dependencies in a row multiply

If a request needs every service on its path, it only succeeds when all of them are up. Multiply their availabilities.

Load balancer99.99%API99.9%Auth99.9%Database99.9%Load balancer99.99%API99.9%Auth99.9%Database99.9%

0.9999 × 0.999 × 0.999 × 0.999 ≈ 99.69%: worse than every single part, about 2.2 hours down a month.

4 · The fix

Copies in parallel add nines

Put two copies side by side and the request fails only when both are down at once: 1 − (1 − a)². Two 99% servers give 99.99%, if their failures are independent.

Load balancerServer A99%Server B99%

Both down at the same time: 1% × 1% = 0.01%. Together: 99.99%.

That if carries the weight. Two servers on the same rack, the same bad deploy, the same expired certificate or the same region fail together. Real redundancy means different racks, zones and release times.

5 · In a real system

When one link has no copy

URL Shortener System Design

The cache is a single point of pressure

The URL shortener's redirects can survive the cache dying only if Cassandra can carry every read on its own. Step through what happens when Redis goes down at peak.

Check yourself

3 questions

1. About how much downtime does 99.99% allow in a 30-day month?
2. A request passes through three services, each 99.9% available. What's the request's availability?
3. Two replicas, each 99% available, run in the same rack. Why might you get far less than 99.99%?

Takeaways

Remember this

  • Each nine is ten times less downtime; past three nines recovery must be automatic.
  • SLI is measured, SLO is the target, SLA is the promise; the gap is your error budget.
  • Serial dependencies multiply availability down; parallel copies add nines back, but only if they fail independently.