Microsoft Azure

Streaming — Event Hubs & Stream Analytics

Ingest millions of events per second, let several consumers read independently, and replay when needed.

Event Hubs is Azure's high-throughput ingestion service: an append-only log that many consumers read at their own pace.

A conveyor belt past several workstations rather than a counter where each item is handed to one person. Everything stays on the belt for a while and each station takes what it needs.

Key Concepts

1
    producers -> [ Event Hub: 4 partitions ] -> consumer group "billing"
                                             -> consumer group "analytics"
2
The difference from Service Bus is the one to state.
    Service Bus  a message is consumed and removed. One consumer wins.
    Event Hubs   events stay for the retention period. Every consumer
                 group keeps its own offset and can re-read.
3
So Event Hubs is the choice when several systems need the same stream, or when replaying after fixing a bug is a requirement.
4
A partition is the unit of ordering and parallelism.
    ordering is guaranteed WITHIN a partition, never across them
    one active reader per partition per consumer group
    so partitions cap your parallelism
5
The partition key decides the partition, and this is where designs fail:
6
    partitionKey = deviceId    events for a device stay ordered,
                               spread across all partitions
    partitionKey = "events"    everything on one partition: a hot
                               partition, and the rest sit idle
7
Throughput is bought in units.
    Throughput Units   1 MB/s in, 2 MB/s out each (Standard)
    Processing Units   Premium
    Capacity Units     Dedicated
8
Auto-inflate raises TUs automatically under load, and does not lower them again — a common bill surprise.
9
Capture writes the stream to storage automatically, in Avro, with no code. That is the cheap way to keep raw events for replay or batch analysis.
10
Consumer groups are how teams stay independent. Each has its own offsets, so a slow analytics consumer cannot delay billing.
11
Stream Analytics queries the stream in SQL.
    SELECT deviceId, AVG(temp)
    INTO   alerts
    FROM   input TIMESTAMP BY eventTime
    GROUP BY deviceId, TumblingWindow(minute, 5)
12
Windowing — tumbling, hopping, sliding, session — is the part worth knowing, along with the distinction between event time and arrival time.
13
It speaks Kafka. Event Hubs exposes a Kafka endpoint, so existing Kafka producers and consumers connect by changing configuration rather than code.
14
What the interviewer is probing.1. "How does Event Hubs differ from Service Bus?" Probing: retention and consumption. Stalls: "It is faster." Moves up: Service Bus removes a consumed message; Event Hubs retains events so every consumer group reads independently and can replay.
15
2. "What limits your parallelism?" Probing: partitions. Stalls: "Throughput units." *Moves up:* one active reader per partition per consumer group, so the partition count caps concurrent consumers — and it is fixed at creation on Standard.
16
3. "Why did your bill rise after enabling auto-inflate?" Probing: the one-way scaling. Stalls: "More traffic." Moves up: auto-inflate raises throughput units under load and never lowers them again.
17
4. "What is the difference between event time and arrival time?" Probing: stream processing correctness. Stalls: "They are the same." Moves up: events arrive late and out of order, so windowing on arrival time gives silently wrong aggregates; Stream Analytics uses TIMESTAMP BY for event time.