Microsoft Azure

Resilience & Disaster Recovery — Zones and Paired Regions

Design for failure against a stated RTO and RPO, and pick a DR strategy the business will actually pay for.

Start with two numbers. An interviewer expects you to ask for them.

Insurance. Comprehensive cover costs more and pays out immediately; basic cover is cheap and leaves you waiting. Nobody can choose for you without knowing what a day of downtime costs.

Key Concepts

1
    RTO   how long may it be down?      recovery time objective
    RPO   how much data may be lost?    recovery point objective
2
"Highly available" without those is not a design.
3
The failure domains, in order.
    instance   replaced automatically by a scale set
    zone       a separate datacentre in the region, independent
               power and cooling. Zone redundancy is cheap.
    region     an entirely separate footprint. Expensive.
4
Availability Zones versus Availability Sets is an Azure-specific question:
5
    Availability Set   spreads VMs across fault and update domains
                       INSIDE one datacentre. Protects against rack
                       and host maintenance only.
    Availability Zone  spreads across SEPARATE datacentres.
                       Protects against a datacentre failure.
6
Zones are the stronger guarantee and the modern default; sets exist for regions without zones.
7
Zone-redundant services do it for you. Zone-redundant storage, zone-redundant Application Gateway and SQL, and VMSS spread across zones — these need configuration, not architecture.
8
Paired regions are Azure's planned-maintenance pairing: updates roll out to one of a pair at a time, and some geo-redundant services replicate there by default. Worth naming, though several newer regions are not paired.
9
The four region strategies, cheapest first.
    backup & restore   restore from backups.   RTO hours, RPO hours
    pilot light        data replicating, compute off. RTO ~tens of
                       minutes
    warm standby       a small live copy, scaled on failover. RTO minutes
    active-active      both regions serving. RTO near zero, hardest
10
Pilot light is the usual answer for a serious application: the expensive part, the data, is already there and only compute must start.
11
Active-active forces a data decision — Cosmos DB multi-region writes or SQL failover groups — and an answer for conflicting writes.
12
A backup you have never restored is an assumption. Test it, time it, and compare that with the RTO you promised. And guard against the realistic disaster, which is a mistaken delete: soft delete, resource locks and point-in-time restore.
13
What the interviewer is probing.1. "What two numbers start a DR design?" Probing: requirements first. Stalls: "Make it redundant." Moves up: RTO and RPO — how long it may be down and how much data may be lost.
14
2. "Availability Set or Availability Zone for a production VM?" Probing: the Azure distinction. Stalls: "Either is fine." Moves up: zones, which are separate datacentres; a set only spreads across racks inside one.
15
3. "What are paired regions?" Probing: the Azure concept. Stalls: "Any two regions." *Moves up:* a designated pair where platform updates roll out to one at a time and some geo-redundant services replicate by default — though several newer regions are unpaired.
16
4. "Describe pilot light." Probing: the strategies. Stalls: "A backup region." Moves up: data replicating continuously with compute switched off, so recovery is starting compute — tens of minutes, far cheaper than active-active.