System design2 min
Message Queues & Pub/Sub
To implement an Event-Driven Architecture or to decouple heavy background processing from your web servers, you need an intermediary system to hold and route the messages.
There are two primary models for this: Message Queues and Publish/Subscribe (Pub/Sub).
1. Message Queues (Point-to-Point)
In a Message Queue model, a message is produced and sent to a specific queue. Multiple consumers might be listening to that queue, but only one consumer will successfully process the message.
Use Case: Task Distribution.
Imagine a video processing app. Users upload videos, and the web server drops a ProcessVideo_ID123 task into the queue. You have 10 Worker Servers listening to the queue. You absolutely do not want all 10 servers to process the same video. The queue ensures that Worker A grabs Video 1, Worker B grabs Video 2, etc.
Key Features:
- Load Leveling: If 1,000 users upload a video at the exact same time, your database and web servers won't crash. The queue absorbs the spike. The workers process the queue at their own sustainable pace.
- Example Tools: RabbitMQ, Amazon SQS, Celery.
2. Publish/Subscribe (Pub/Sub)
In a Pub/Sub model, a message is published to a "Topic." Multiple consumers (Subscribers) can listen to that topic, and every subscriber receives a copy of the message.
Use Case: Event Broadcasting.
When a user buys a product, the Checkout Service publishes an OrderPlaced event to the topic. The Inventory Service receives it and updates stock. The Shipping Service receives it and prints a label. The Analytics Service receives it and updates the sales dashboard.
Key Features:
- Fan-out: One message triggers multiple independent parallel workflows in different services.
- Example Tools: Apache Kafka, Amazon SNS, Google Cloud Pub/Sub.
Apache Kafka: The Heavyweight
Kafka often comes up in System Design interviews. It is fundamentally different from a traditional queue like RabbitMQ.
Traditional queues delete a message as soon as it is successfully processed. Kafka, however, is a Distributed Commit Log. It stores messages permanently on disk (for a configurable retention period, e.g., 7 days).
Why does this matter? If you deploy a bug in your Analytics Service that corrupts your data for 3 days, with RabbitMQ, those events are gone forever. With Kafka, you can simply "rewind" your Analytics Service's pointer to 3 days ago, and it will re-process every historical event, fixing your data!