Microsoft Azure
Resilience & Disaster Recovery — Zones and Paired Regions
Design for failure against a stated RTO and RPO, and pick a DR strategy the business will actually pay for.
Start with two numbers. An interviewer expects you to ask for them.
Insurance. Comprehensive cover costs more and pays out immediately; basic cover is cheap and leaves you waiting. Nobody can choose for you without knowing what a day of downtime costs.
Key Concepts
1
RTO how long may it be down? recovery time objective
RPO how much data may be lost? recovery point objective2
"Highly available" without those is not a design.
3
The failure domains, in order.
instance replaced automatically by a scale set
zone a separate datacentre in the region, independent
power and cooling. Zone redundancy is cheap.
region an entirely separate footprint. Expensive.4
Availability Zones versus Availability Sets is an Azure-specific question:
5
Availability Set spreads VMs across fault and update domains
INSIDE one datacentre. Protects against rack
and host maintenance only.
Availability Zone spreads across SEPARATE datacentres.
Protects against a datacentre failure.6
Zones are the stronger guarantee and the modern default; sets exist for regions without zones.
7
Zone-redundant services do it for you. Zone-redundant storage, zone-redundant Application Gateway and SQL, and VMSS spread across zones — these need configuration, not architecture.
8
Paired regions are Azure's planned-maintenance pairing: updates roll out to one of a pair at a time, and some geo-redundant services replicate there by default. Worth naming, though several newer regions are not paired.
9
The four region strategies, cheapest first.
backup & restore restore from backups. RTO hours, RPO hours
pilot light data replicating, compute off. RTO ~tens of
minutes
warm standby a small live copy, scaled on failover. RTO minutes
active-active both regions serving. RTO near zero, hardest10
Pilot light is the usual answer for a serious application: the expensive part, the data, is already there and only compute must start.
11
Active-active forces a data decision — Cosmos DB multi-region writes or SQL failover groups — and an answer for conflicting writes.
12
A backup you have never restored is an assumption. Test it, time it, and compare that with the RTO you promised. And guard against the realistic disaster, which is a mistaken delete: soft delete, resource locks and point-in-time restore.
13
What the interviewer is probing.1. "What two numbers start a DR design?" Probing: requirements first. Stalls: "Make it
redundant." Moves up: RTO and RPO — how long it may be down and how much data may be lost.
14
2. "Availability Set or Availability Zone for a production VM?" Probing: the Azure
distinction. Stalls: "Either is fine." Moves up: zones, which are separate datacentres; a set
only spreads across racks inside one.
15
3. "What are paired regions?" Probing: the Azure concept. Stalls: "Any two regions." *Moves
up:* a designated pair where platform updates roll out to one at a time and some geo-redundant
services replicate by default — though several newer regions are unpaired.
16
4. "Describe pilot light." Probing: the strategies. Stalls: "A backup region." Moves up:
data replicating continuously with compute switched off, so recovery is starting compute — tens of
minutes, far cheaper than active-active.