System design3 min
Disaster Recovery & Redundancy
A system is not truly highly available unless it can survive a catastrophic event—like a data center catching fire, a massive power grid failure, or a severed undersea cable.
Disaster Recovery (DR) is the strategy for returning an application to a functional state after such an event.
RPO and RTO
Every DR plan is defined by two critical business metrics:
1. Recovery Point Objective (RPO)
"How much data can we afford to lose?" If your RPO is 1 hour, it means you must back up your data every hour. If the server explodes, you will lose, at most, the last 59 minutes of user data. For a social media site, an RPO of 1 hour might be acceptable. For a stock exchange, the RPO must be zero milliseconds.
2. Recovery Time Objective (RTO)
"How long can the application be offline?" If your RTO is 4 hours, it means your engineering team has 4 hours to provision new servers, restore the database from backups, update DNS, and get the site back online. For critical infrastructure, RTO is often measured in seconds.
Disaster Recovery Strategies
These strategies are often categorized by "temperature" (how active the secondary systems are).
1. Cold Standby (Backup & Restore)
- How it works: You take daily snapshots of your database and store them securely in a different geographic region (e.g., AWS S3).
- In a disaster: You manually spin up new servers in the new region, download the massive backup file, restore the database, and update DNS.
- Metrics: High RTO (takes hours/days to restore), High RPO (you lose up to 24 hours of data).
- Cost: Very cheap.
2. Warm Standby (Pilot Light)
- How it works: You maintain a live, constantly replicating database in a secondary region. However, your web servers and application logic in the secondary region are either scaled down to zero or very minimal.
- In a disaster: The database is already there and up to date. You just need to run your deployment scripts to rapidly spin up the web servers and switch the DNS.
- Metrics: Medium RTO (minutes to hours), Low RPO (seconds, due to replication).
- Cost: Moderate.
3. Hot Standby (Active-Passive)
- How it works: You have a fully provisioned, 100% identical infrastructure running in the secondary region. It receives real-time database replication, but no active user traffic.
- In a disaster: You literally just flip a switch in your DNS provider to point users from Region A to Region B.
- Metrics: Very Low RTO (seconds), Low RPO (seconds).
- Cost: Expensive (you are paying for an entire duplicate data center that does nothing 99% of the time).
4. Multi-Region Active-Active
- How it works: You run fully functional infrastructure in multiple regions simultaneously (e.g., US-East and EU-West). Both regions accept read and write traffic.
- In a disaster: If US-East goes offline, the global load balancer automatically routes all US users to EU-West.
- Metrics: Near-Zero RTO, Near-Zero RPO.
- Cost: Extremely expensive and highly complex (requires multi-master database replication and complex conflict resolution).