System designCore3 min
Availability, SLAs & nines
What 99.99% available really allows, why a chain of services is less available than any link in it, and how redundancy wins it back.
1 · Counting nines
Each nine is ten times less downtime
Availability is the share of time, or of requests, that the system serves correctly. It's quoted in nines, and each extra nine cuts the allowed downtime by ten.
| Availability | Down per year | Down per 30 days |
|---|---|---|
| 99% (two nines) | 3.65 days | 7.2 hours |
| 99.9% (three nines) | 8.8 hours | 43 minutes |
| 99.95% | 4.4 hours | 22 minutes |
| 99.99% (four nines) | 53 minutes | 4.3 minutes |
| 99.999% (five nines) | 5.3 minutes | 26 seconds |
Four nines leaves about four minutes a month: not enough time for a person to be paged, log in and fix anything. Past three nines, recovery has to be automatic.
2 · The vocabulary
SLI, SLO, SLA and the error budget
| Term | What it is | Example |
|---|---|---|
| SLI (indicator) | What you measure | Share of requests that succeed in under 300 ms |
| SLO (objective) | The target you hold yourself to | 99.9% of requests over any 30 days |
| SLA (agreement) | A promise to customers, with a penalty | 99.5%, or a service credit. Set looser than the SLO |
| Error budget | What the SLO lets you miss | 0.1% of requests: spend it on releases and experiments |
The error budget turns reliability into a decision: while there's budget left, ship; when it's spent, slow down and fix things.
3 · The catch
Dependencies in a row multiply
If a request needs every service on its path, it only succeeds when all of them are up. Multiply their availabilities.
0.9999 × 0.999 × 0.999 × 0.999 ≈ 99.69%: worse than every single part, about 2.2 hours down a month.
4 · The fix
Copies in parallel add nines
Put two copies side by side and the request fails only when both are down at once: 1 − (1 − a)². Two 99% servers give 99.99%, if their failures are independent.
Both down at the same time: 1% × 1% = 0.01%. Together: 99.99%.
That if carries the weight. Two servers on the same rack, the same bad deploy, the same expired certificate or the same region fail together. Real redundancy means different racks, zones and release times.
5 · In a real system
When one link has no copy
URL Shortener System Design
The cache is a single point of pressure
The URL shortener's redirects can survive the cache dying only if Cassandra can carry every read on its own. Step through what happens when Redis goes down at peak.
Check yourself
3 questions
Takeaways
Remember this
- Each nine is ten times less downtime; past three nines recovery must be automatic.
- SLI is measured, SLO is the target, SLA is the promise; the gap is your error budget.
- Serial dependencies multiply availability down; parallel copies add nines back, but only if they fail independently.