Amazon Web Services
Resilience & Disaster Recovery — Multi-AZ and Multi-Region
Design for failure with a stated RTO and RPO, and pick a DR strategy that matches what the business will pay for.
Resilience starts with two numbers, and an interviewer will expect you to ask for them.
Insurance. Comprehensive cover costs more than you want but pays out immediately; basic cover is cheap and leaves you waiting. Nobody can choose for you without knowing what a day without the car costs.
Key Concepts
1
RTO how long may it be down? -> recovery time objective
RPO how much data may be lost? -> recovery point objective2
Everything else follows. "Highly available" without those numbers is not a design.
3
The hierarchy of failure domains.
instance replace it -- Auto Scaling does this already
AZ separate buildings, independent power.
Multi-AZ is the baseline and is cheap.
region an entirely separate footprint. Expensive.4
Multi-AZ is the default and is not DR. RDS Multi-AZ keeps a synchronous standby in another AZ and fails
over in a minute or two. It protects against an AZ loss, not a bad deploy or a deleted table.
5
The four region strategies, cheapest first.
backup & restore restore from snapshots. RTO hours, RPO hours.
pilot light data replicating, compute off. RTO ~tens of
minutes, RPO minutes.
warm standby a small live copy, scaled up on failover.
RTO minutes.
active-active both regions serving. RTO near zero, and the
hardest and most costly to run.6
Pilot light is the usual answer for a serious application, because the expensive part — the data — is
already there and only the compute must start.
7
Active-active forces a data decision. Two regions taking writes means either DynamoDB global tables or
Aurora Global Database, and an answer for conflicting writes. Most teams discover they do not need it.
8
A backup you have never restored is not a backup. Test the restore, measure how long it took, and
compare that with the RTO you promised.
9
Guard against the non-infrastructure failures too. Versioning and MFA delete on S3, point-in-time
recovery on DynamoDB and RDS, and deletion protection — because the realistic disaster is a mistaken
DELETE, not a region falling over.
DELETE
10
What the interviewer is probing.1. "What do you ask before designing for disaster recovery?" Probing: whether you start from
requirements. Stalls: "Make it highly available." Moves up: the RTO and RPO — how long it may be
down and how much data may be lost. Everything else follows from those two numbers.
11
2. "Does Multi-AZ protect you from a bad deploy?" Probing: what redundancy does not cover.
Stalls: "Yes, it is redundant." Moves up: no — it covers infrastructure failure, not a bad
release or a deleted table; that needs versioning, point-in-time recovery and deletion protection.
12
3. "Explain pilot light." Probing: the strategies. Stalls: "A standby environment." *Moves
up:* the data is already replicating to the second region while compute stays off, so recovery is
starting compute — tens of minutes, at a fraction of active-active cost.
13
4. "What makes active-active hard?" Probing: the data problem. Stalls: "The cost." *Moves
up:* two regions taking writes means conflict resolution — global tables or an Aurora Global
Database — and most teams find they do not need it.