Google Cloud

Resilience & Disaster Recovery — Zones, Regions & Multi-region

Design for failure against a stated RTO and RPO, and use the global network rather than DNS to fail over.

Start with two numbers, and expect to be asked for them.

Insurance. Comprehensive cover costs more and pays out immediately; basic cover is cheap and leaves you waiting. Nobody can choose without knowing what a day of downtime costs.

Key Concepts

1
    RTO   how long may it be down?      recovery time objective
    RPO   how much data may be lost?    recovery point objective
2
"Highly available" without those is not a design.
3
The failure domains.
    instance   a MIG auto-heals it
    zone       a separate datacentre in the region
    region     an entirely separate footprint
    multi-region  a Google concept: storage and services spanning
                  a continent
4
Regional resources are the baseline. A regional MIG spreads instances across zones automatically, a regional Persistent Disk replicates synchronously across two, and regional Cloud Storage keeps data in several. Most of resilience here is choosing "regional" instead of "zonal" at creation.
5
GCP's global load balancer changes the failover story. One anycast IP worldwide means a failed region is routed around by the network in seconds — no DNS change, no TTL to wait out. That is a genuine architectural advantage and worth stating explicitly.
6
The strategies, cheapest first.
    backup & restore   restore from snapshots.  RTO hours
    pilot light        data replicating, compute off. RTO ~tens of
                       minutes
    warm standby       a small live deployment, scaled on failover
    active-active      both regions serving behind one IP
7
Active-active is easier here than elsewhere, because the global load balancer already fronts both regions. The hard part remains the data.
8
The data decision is what makes it real.
    Spanner        multi-region, strongly consistent, external
                   consistency via TrueTime. The strongest option.
    Cloud SQL      cross-region read replicas, manual promotion
    Firestore      multi-region replication built in
    Cloud Storage  dual-region or multi-region buckets
9
Spanner is the differentiator when an application genuinely needs relational semantics and multi-region writes without conflict resolution.
10
Backups you have never restored are assumptions. Test the restore, time it, compare with the RTO you promised — and guard against the realistic disaster, a mistaken delete, with object versioning, retention policies and deletion protection.
11
What the interviewer is probing.1. "What two numbers start the design?" Probing: requirements first. Stalls: "Make it redundant." Moves up: RTO and RPO — how long it may be down and how much data may be lost.
12
2. "What makes active-active easier on GCP than elsewhere?" Probing: the anycast front door. Stalls: "Nothing." Moves up: the global load balancer already fronts both regions behind one IP, so the compute side is largely solved — the data side is where the work remains.
13
3. "What is the most common resilience mistake?" Probing: the default. Stalls: "Not enough replicas." Moves up: accepting zonal resources — a zonal MIG, disk or cluster — when the regional equivalent costs the same and removes the failure mode.
14
4. "When is Spanner the right DR answer?" Probing: proportionality. Stalls: "Always, it is multi-region." Moves up: when the application needs relational semantics with multi-region writes and no conflict resolution — otherwise it is expensive for what it gives.