Google Cloud
Resilience & Disaster Recovery — Zones, Regions & Multi-region
Design for failure against a stated RTO and RPO, and use the global network rather than DNS to fail over.
Start with two numbers, and expect to be asked for them.
Insurance. Comprehensive cover costs more and pays out immediately; basic cover is cheap and leaves you waiting. Nobody can choose without knowing what a day of downtime costs.
Key Concepts
1
RTO how long may it be down? recovery time objective
RPO how much data may be lost? recovery point objective2
"Highly available" without those is not a design.
3
The failure domains.
instance a MIG auto-heals it
zone a separate datacentre in the region
region an entirely separate footprint
multi-region a Google concept: storage and services spanning
a continent4
Regional resources are the baseline. A regional MIG spreads instances across zones automatically, a regional Persistent Disk replicates synchronously across two, and regional Cloud Storage keeps data in several. Most of resilience here is choosing "regional" instead of "zonal" at creation.
5
GCP's global load balancer changes the failover story. One anycast IP worldwide means a failed region is routed around by the network in seconds — no DNS change, no TTL to wait out. That is a genuine architectural advantage and worth stating explicitly.
6
The strategies, cheapest first.
backup & restore restore from snapshots. RTO hours
pilot light data replicating, compute off. RTO ~tens of
minutes
warm standby a small live deployment, scaled on failover
active-active both regions serving behind one IP7
Active-active is easier here than elsewhere, because the global load balancer already fronts both regions. The hard part remains the data.
8
The data decision is what makes it real.
Spanner multi-region, strongly consistent, external
consistency via TrueTime. The strongest option.
Cloud SQL cross-region read replicas, manual promotion
Firestore multi-region replication built in
Cloud Storage dual-region or multi-region buckets9
Spanner is the differentiator when an application genuinely needs relational semantics and multi-region writes without conflict resolution.
10
Backups you have never restored are assumptions. Test the restore, time it, compare with the RTO you promised — and guard against the realistic disaster, a mistaken delete, with object versioning, retention policies and deletion protection.
11
What the interviewer is probing.1. "What two numbers start the design?" Probing: requirements first. Stalls: "Make it
redundant." Moves up: RTO and RPO — how long it may be down and how much data may be lost.
12
2. "What makes active-active easier on GCP than elsewhere?" Probing: the anycast front door.
Stalls: "Nothing." Moves up: the global load balancer already fronts both regions behind one IP,
so the compute side is largely solved — the data side is where the work remains.
13
3. "What is the most common resilience mistake?" Probing: the default. Stalls: "Not enough
replicas." Moves up: accepting zonal resources — a zonal MIG, disk or cluster — when the regional
equivalent costs the same and removes the failure mode.
14
4. "When is Spanner the right DR answer?" Probing: proportionality. Stalls: "Always, it is
multi-region." Moves up: when the application needs relational semantics with multi-region writes
and no conflict resolution — otherwise it is expensive for what it gives.