Google Cloud

Streaming & Analytics — Dataflow & BigQuery

Process streams and batches with one programming model, and query petabytes without managing a cluster.

GCP's data story is its strongest area, and two services carry it.

A sorting line that groups parcels by when they were posted rather than when they arrived, and holds each batch open just long enough for the stragglers.

Key Concepts

1
    Dataflow   managed Apache Beam. ONE pipeline for streaming
               and batch, autoscaling, no cluster to size.
    BigQuery   serverless warehouse. SQL over petabytes, storage
               and compute billed separately.
2
Beam's unified model is the distinguishing idea. The same transforms run over a bounded source (batch) or an unbounded one (streaming), so you do not maintain two implementations of the same logic.
3
Windowing is where streaming gets real.
    fixed     every 5 minutes, non-overlapping
    sliding   5-minute window every 1 minute, overlapping
    session   group by gaps in activity per user
4
Event time versus processing time is the concept interviewers probe. Events arrive late and out of order; Beam processes by event time and uses a watermark to estimate how far along it is, with triggers deciding when to emit and how to handle stragglers.
5
That is why a count for 10:00–10:05 can be emitted, then updated when a late event arrives — and why you must decide whether to accumulate or discard.
6
BigQuery separates storage from compute. You are billed for bytes scanned, which makes two habits essential:
7
    SELECT *                 scans every column -- expensive
    SELECT order_id, total   scans two
8
    partition by date, cluster by customer_id
      -> a WHERE on the partition column skips the rest entirely
9
Partitioning and clustering are the cost control. An unpartitioned table means every query reads everything.
10
Streaming into BigQuery through the Storage Write API makes rows queryable within seconds, which is what makes "near real-time dashboard" straightforward here.
11
The common pipeline.
    Pub/Sub -> Dataflow (window, aggregate, enrich) -> BigQuery
                      \-> Cloud Storage (raw, for replay)
12
Dataflow autoscales workers on backlog and CPU, and drains cleanly on update so in-flight work completes rather than being dropped.
13
What the interviewer is probing.1. "What is the Beam model's main idea?" Probing: unification. Stalls: "It processes data." Moves up: the same transforms run over bounded and unbounded sources, so batch and streaming are one pipeline rather than two implementations.
14
2. "Explain watermarks and late data." Probing: the hard part of streaming. Stalls: "Events arrive in order." Moves up: processing is by event time and the watermark estimates progress; a late event can re-trigger a window, and allowedLateness decides whether it is counted or dropped.
15
3. "Why did one BigQuery query cost so much?" Probing: the billing model. Stalls: "It scanned a lot of rows." Moves up: billing is by bytes scanned, so SELECT * on a wide unpartitioned table reads every column of every partition.
16
4. "How do you control BigQuery cost structurally?" Probing: the table design. Stalls: "Tell people to be careful." Moves up: partition by date and cluster by a common filter column, select only needed columns, and set custom quotas so one query cannot run away.