System designHard4 min readruns in the simulator
Netflix System Design (High-Scale)
Open Connect CDN for playback, a queue-driven transcode pipeline: flush the cache and watch the origin take the hit.
- cdn
- streaming
- storage
- microservices
- queue
- throughput
- 6,000 rps
- play p99 latency
- 134 ms
- cost
- $10,417/mo
Requirements
Functional
- Users can browse a personalized home feed of recommended titles.
Non-functional
- Playback starts within 500 ms p99 for content already on the CDN edge.play p99 latency < 500 ms
- Support 200M subscribers, up to 50M concurrent streams at peak (modeled here at a scaled-down 6,000 req/s so the simulation runs at a workable size).play successful requests ≥ 4000 rps
The requirements with a measurable target are checked live in the simulator: break something and watch them fail.
High-level architecture
Smart TV Appclient · tv
Mobile Appclient · mobile
Zuul API GatewayapiGateway
Open Connect CDNcdn
Auth Serviceserver · api
Small-scale compute cluster
Billing Serviceserver · api
Small-scale compute cluster
Geo-Licensing Serviceserver · api
Small-scale compute cluster
Recommendation Engineserver
Transcode Schedulerorchestrator
Transcode Task Queuequeue
GPU Transcoder Nodeworker
Video Block Store (S3)objectStore
Billing DB (MySQL)db
User Profiles DB (Cassandra)db · cassandra
Distributed NoSQL database with 3 replicas
Session & Geo Cachecache · redis
Single-shard in-memory cache
Telemetry Bus (Kafka)messageBus · kafka
Distributed message broker
MaxMind Geo-IP Servicecloud
Studio Ingest Clientclient · studio
Request paths
Browse the home feed
Zuul API GatewayRecommendation EngineUser Profiles DB (Cassandra)
Start playback
Open Connect CDNon a missVideo Block Store (S3)
Start playback via the CDN
- Press play.
70% of viewer requests are playback. The TV and mobile apps fetch video segments straight from the Open Connect edge; the API gateway isn't on this path.
- Served from the edge.
92% of segment requests are edge hits, served in about 10 ms.
- A miss goes to S3.
The rest are read through from the S3 block store, about 60 ms more, and kept at the edge for the next viewer. A cold edge takes around 45 s to warm up.
- The target.
Playback must start within 500 ms p99 for content already at the edge, and hits keep it far inside that.
Studio upload & encode
Video Block Store (S3)Transcode SchedulerTranscode Task Queue
Video transcoding pipeline
- Upload the master.
A studio uploads the master file to the S3 block store.
- Schedule the encode.
The upload triggers the transcode scheduler, which splits the work into encoding tasks.
- Queue the tasks.
Tasks wait in the queue, up to 5,000 deep, so a burst of uploads never swamps the encoders.
- Encode on GPUs.
10 GPU workers pull tasks, about 5 a second each, and write the encoded blocks back to S3, ready for the CDN.
What happens when it breaks
The edge goes dark
Open Connect fails entirely: every playback start errors out, but browsing the catalog keeps working — the edge is a single point of failure for playback only.
| Reading | Healthy | During the failure |
|---|---|---|
| p99 latency | 252 ms | 251 ms |
| error rate | 0% | 70% |
| throughput | 6,000 rps | 1,800 rps |
| cost | $10,417/mo | $10,417/mo |
- Playback starts within 500 ms p99 for content already on the CDN edge. (fails)
- Support 200M subscribers, up to 50M concurrent streams at peak (modeled here at a scaled-down 6,000 req/s so the simulation runs at a workable size). (fails)
Trade-offs
Fixes and what they cost
| Fix | What it does | Trade-off |
|---|---|---|
| Coalesce CDN misses | Concurrent misses for the same title share one origin fetch, which prevents a thundering herd on the object store. | Viewers asking for the same title wait for the first origin fetch. |
| Add origin read capacity | More origin read throughput, at a real cost per month. | Costs more per month for origin capacity that sits idle most of the time. |
| Pre-warm the CDN edge | A flushed or restarted edge loads 90% of the popular titles before it serves, so the origin never sees a cold-cache stampede. It takes effect on the next flush or restart. | A restarted edge takes longer to come back while it loads popular titles. |
| Circuit breaker on gateway to recommendations | When recommendations fail, the gateway stops waiting on them and serves a generic row. Latency drops, and the struggling service gets room to recover. | Viewers see a generic row instead of personal recommendations while it's open. |