System design2 min
Event-Driven Architecture
In traditional request-response architectures (like REST or RPC), services communicate synchronously. If Service A needs Service B to do something, it calls Service B and waits for an answer.
This creates tight coupling. If Service B is down, Service A fails. If Service A needs to notify Service C and Service D as well, Service A has to be updated to know about them, and it has to wait for all of them to finish.
Event-Driven Architecture (EDA) solves this by changing the communication paradigm from "commands" to "events."
How it Works
Instead of Service A telling Service B to do something, Service A simply broadcasts an Event stating that something happened in the past (e.g., UserCreated, OrderPlaced, PaymentFailed).
Service A does not care who is listening. It drops the event into an Event Bus (like Kafka or RabbitMQ) and goes back to work.
Other services (Subscribers) listen to the Event Bus. When the OrderPlaced event arrives, the Inventory Service catches it and reserves stock. The Notification Service catches it and emails the user. The Analytics Service catches it and updates the dashboard.
Pros of EDA
- Loose Coupling: The Order Service doesn't know the Notification Service exists. You can add a completely new Fraud Detection Service tomorrow that listens to
OrderPlaced, and you won't have to change a single line of code in the Order Service. - Fault Tolerance: If the Notification Service crashes, the Order Service isn't affected. The
OrderPlacedevent just sits in the queue. When the Notification Service reboots, it processes the backlog. - Scalability: You can easily spin up multiple instances of the Notification Service to process events in parallel if the queue gets too long.
Cons of EDA
- Eventual Consistency: Because events are processed asynchronously in the background, there is a delay. A user might place an order, refresh their dashboard instantly, and not see the order yet because the database update event is still sitting in the queue.
- Complexity in Tracking: If an order fails, it's very hard to figure out why. You have to trace the flow of events across 5 different services to find the bug. Distributed tracing tools (like Jaeger or Zipkin) become mandatory.
- The "Two Generals" Problem (Idempotency): Networks are unreliable. What happens if the Event Bus crashes and accidentally delivers the
ChargeCreditCardevent twice? Subscribers must be designed to be idempotent, meaning they can safely process the exact same event multiple times without causing side effects.