Apache Cassandra

Materialized Views in Cassandra

Learn how materialized views automate denormalization for alternate query patterns, and understand their operational risks.

Query-first modelling means writing the same order into orders_by_customer and orders_by_product yourself. A materialized view was meant to do that for you.

A materialized view is like having an assistant who automatically re-files a copy of every document you create into a second cabinet organized differently — convenient, but if the assistant ever misses an update during a busy day, the two cabinets quietly stop matching.

Key Concepts

1
You define the second table once, Cassandra keeps it filled.
2
    CREATE MATERIALIZED VIEW orders_by_product AS
      SELECT * FROM orders
      WHERE product_id IS NOT NULL AND order_id IS NOT NULL
      PRIMARY KEY ((product_id), order_date, order_id);
3
Write to orders and the view updates. No application code, no forgotten table.
orders
4
Why it is harder than it looks. The base row and the view row usually live on different nodes, because they have different partition keys and therefore different tokens.
5
So the node holding the base row must also update a node somewhere else, and keep doing so correctly while nodes restart, repairs run, and data streams during scaling. That is a distributed transaction between two partitions — the thing Cassandra deliberately does not do.
6
What happens in practice. The view drifts. A row exists in the base table and is missing from the view, or the view holds a row the base table no longer has.
7
Nothing reports this. There is no consistency check, and nodetool repair repairs the base table and the view separately, so it can leave them disagreeing. Teams discover it when a query returns results a user knows are wrong.
nodetool repair
8
So they were marked experimental, with a warning printed at creation, and the Cassandra project itself advises against production use. Not removed, because existing clusters depend on them.
9
What to do instead. Write both tables from the application, in one place in the codebase. That is more code and it fails in ways you can see, test and fix.
10
If the two writes must both land, use a logged batch — atomicity is exactly what the batchlog gives you.
11
Why interviewers ask. The question is rarely about the syntax. It is whether you can explain why a feature that removes obvious boilerplate turned out to be a bad trade, which is a better test of judgement than knowing the feature exists.
12
What the interviewer is probing.1. "What problem do they solve?" Probing: the purpose. Stalls: "Caching." Moves up: maintaining a second table keyed differently, so a query by another column stays partition-local — without the application writing both.
13
2. "What is their status?" Probing: the honest answer. Stalls: "Production ready." *Moves up:* marked experimental with known consistency edge cases, so many teams maintain the second table in application code instead.
14
3. "What is the constraint on the view key?" Probing: the rule. Stalls: "None." Moves up: it must include every column of the base primary key, so a view cannot aggregate or drop key columns.
15
4. "What is the write cost?" Probing: the trade. Stalls: "None." Moves up: each base write triggers a read and a write on the view, which amplifies write load.