Microsoft Azure
Cosmos DB Data Modelling — Partition Keys & RU/s
Choose a partition key for the access patterns and understand request units, or the bill and the throttling will tell you.
Cosmos DB is fast and predictable only if the partition key matches the queries. This is the most-asked Azure database topic.
A filing room where the partition key is the cabinet. Put everything under one label and there is a queue at that cabinet, no matter how many others stand empty.
Key Concepts
1
Everything is priced in Request Units.
point read of a 1 KB item ~1 RU
a simple query a few RU
a write ~5 RU and up
a cross-partition query RU proportional to partitions touched2
Exceed the provisioned RU/s and you get 429 Too Many Requests. The SDK retries automatically, which is why the symptom is usually latency rather than errors.
429 Too Many Requests
3
The partition key decides everything.
good high cardinality, evenly accessed, and present in
most queries -- customerId, tenantId, deviceId
bad low cardinality or time-based -- "country", or today's
date, which makes every write hit one partition4
A logical partition is capped at 20 GB, and that is a hard limit. A key that puts too much under one value eventually cannot grow, and the fix is a new container with a different key.
5
A cross-partition query is the common performance bug.
SELECT * FROM c WHERE c.email = 'x@y.com' -- fans out
SELECT * FROM c WHERE c.customerId = '123' -- one partition6
Without the partition key in the filter, the query touches every physical partition and the RU cost scales with them.
7
Throughput modes.
provisioned fixed RU/s. Cheapest when steady.
autoscale scales between 10% and the maximum you set.
serverless pay per request. Spiky or small workloads.8
The five consistency levels are Cosmos's distinguishing feature, and session is the default:
9
strong bounded staleness session consistent prefix eventual10
Weaker consistency costs fewer RUs and gives lower latency. Session consistency — read your own writes — is what most applications actually need.
11
Denormalise deliberately. There are no joins across containers, so embed what is read together and accept duplication, keeping it consistent via the change feed.
12
The change feed emits every insert and update in partition order, which is how you maintain materialised views and trigger downstream work.
13
What the interviewer is probing.1. "How do you choose a partition key?" Probing: the central skill. Stalls: "Something
unique like an id." Moves up: high cardinality, evenly accessed, and present in most queries — a
unique id gives perfect distribution but makes every query cross-partition.
14
2. "What is the hard limit on a logical partition?" Probing: the cap people hit late.
Stalls: "There is none." Moves up: 20GB — a key that puts too much under one value eventually
cannot grow, and the fix is a new container.
15
3. "Your queries are expensive in RUs. What is the likely cause?" Probing: cross-partition
fan-out. Stalls: "The data is large." Moves up: the filter omits the partition key, so the query
touches every physical partition and costs scale with them.
16
4. "Can you change the partition key later?" Probing: the migration cost. Stalls: "Yes, in
the portal." Moves up: no — it means a new container and a data migration, which is why the
decision deserves time up front.