System design2 min
The Circuit Breaker Pattern
In a distributed microservices architecture, network failures and timeouts are inevitable. If Service A depends on Service B, and Service B experiences a sudden lag spike (e.g., a bad database query), Service A's HTTP requests will start timing out.
The Cascading Failure Problem
If Service A has a 5-second timeout, every single request hitting Service A will sit idle for 5 seconds waiting for Service B. Very quickly, all of Service A's threads/connections will be exhausted, and Service A will crash.
But it gets worse: if the API Gateway depends on Service A, it will now start timing out, exhausting its own threads. Soon, the entire application goes offline because one single query in one microservice was slow. This is called a Cascading Failure.
The Circuit Breaker Solution
The Circuit Breaker pattern (inspired by electrical engineering) prevents this by wrapping the HTTP call in a protective proxy that monitors for failures.
It has three states:
1. CLOSED (Normal Operation)
When everything is healthy, the circuit is "Closed" (electricity flows). The Circuit Breaker simply passes requests from Service A to Service B. It monitors the responses. If the failure rate (e.g., timeouts or 500 errors) crosses a certain threshold (say, 50% failures over 10 seconds), it trips the circuit.
2. OPEN (Failing)
The circuit is now "Open" (electricity is blocked). Service A's requests are immediately rejected by the Circuit Breaker. It doesn't even attempt to contact Service B.
Why is this brilliant?
- Service A instantly gets an error back, meaning its threads aren't blocked waiting for a timeout. Service A survives!
- Service B, which is currently struggling under load, is given a moment to "breathe" and recover without being hammered by retries.
3. HALF-OPEN (Testing the Waters)
After a timeout period (e.g., 30 seconds), the Circuit Breaker transitions to "Half-Open". It allows a limited number of test requests to pass through to Service B.
- If the test requests succeed, it assumes Service B has recovered and transitions back to CLOSED.
- If the test requests fail, it assumes Service B is still down and immediately snaps back to OPEN, resetting the timeout timer.
Fallback Mechanisms
When the circuit is OPEN, Service A needs to handle the failure gracefully. Instead of showing an ugly error to the user, it can provide a fallback:
- Cached Data: Return stale data from Redis.
- Degraded Experience: If the "Recommendation Service" is down, just show the user a static "Top 10 Global Hits" list instead of crashing the homepage.
- Queuing: Return a "We are processing this in the background" message and drop the task into a Kafka queue for later.