Amazon Web Services

Resilience & Disaster Recovery — Multi-AZ and Multi-Region

Design for failure with a stated RTO and RPO, and pick a DR strategy that matches what the business will pay for.

Resilience starts with two numbers, and an interviewer will expect you to ask for them.

Insurance. Comprehensive cover costs more than you want but pays out immediately; basic cover is cheap and leaves you waiting. Nobody can choose for you without knowing what a day without the car costs.

Key Concepts

1
    RTO   how long may it be down?      -> recovery time objective
    RPO   how much data may be lost?    -> recovery point objective
2
Everything else follows. "Highly available" without those numbers is not a design.
3
The hierarchy of failure domains.
    instance   replace it -- Auto Scaling does this already
    AZ         separate buildings, independent power.
               Multi-AZ is the baseline and is cheap.
    region     an entirely separate footprint. Expensive.
4
Multi-AZ is the default and is not DR. RDS Multi-AZ keeps a synchronous standby in another AZ and fails over in a minute or two. It protects against an AZ loss, not a bad deploy or a deleted table.
5
The four region strategies, cheapest first.
    backup & restore   restore from snapshots.  RTO hours, RPO hours.
    pilot light        data replicating, compute off. RTO ~tens of
                       minutes, RPO minutes.
    warm standby       a small live copy, scaled up on failover.
                       RTO minutes.
    active-active      both regions serving. RTO near zero, and the
                       hardest and most costly to run.
6
Pilot light is the usual answer for a serious application, because the expensive part — the data — is already there and only the compute must start.
7
Active-active forces a data decision. Two regions taking writes means either DynamoDB global tables or Aurora Global Database, and an answer for conflicting writes. Most teams discover they do not need it.
8
A backup you have never restored is not a backup. Test the restore, measure how long it took, and compare that with the RTO you promised.
9
Guard against the non-infrastructure failures too. Versioning and MFA delete on S3, point-in-time recovery on DynamoDB and RDS, and deletion protection — because the realistic disaster is a mistaken DELETE, not a region falling over.
DELETE
10
What the interviewer is probing.1. "What do you ask before designing for disaster recovery?" Probing: whether you start from requirements. Stalls: "Make it highly available." Moves up: the RTO and RPO — how long it may be down and how much data may be lost. Everything else follows from those two numbers.
11
2. "Does Multi-AZ protect you from a bad deploy?" Probing: what redundancy does not cover. Stalls: "Yes, it is redundant." Moves up: no — it covers infrastructure failure, not a bad release or a deleted table; that needs versioning, point-in-time recovery and deletion protection.
12
3. "Explain pilot light." Probing: the strategies. Stalls: "A standby environment." *Moves up:* the data is already replicating to the second region while compute stays off, so recovery is starting compute — tens of minutes, at a fraction of active-active cost.
13
4. "What makes active-active hard?" Probing: the data problem. Stalls: "The cost." *Moves up:* two regions taking writes means conflict resolution — global tables or an Aurora Global Database — and most teams find they do not need it.