The Thundering Herd: Designing a Live Stream for the Two Million People Who Press Play at the Same Second
Most scaling problems arrive gradually. You watch a graph climb over weeks, you add capacity, you move on.
Live events don't work like that. The broadcast starts at a wall-clock time that millions of people already have in their calendar, and the load curve looks less like a hill than a cliff face. Everyone arrives inside the same ninety seconds, everyone asks for the same bytes, and everyone finds out simultaneously whether you got the design right.
The counterintuitive part — the part I wish someone had told me earlier — is that the video is rarely what breaks. A well-shielded CDN will shrug off ten terabits per second of segment traffic. What breaks is everything that has to make a decision about each of those viewers: is this person allowed to watch, how many devices do they already have open, do they have a valid license.
So let's walk through the whole shape, then spend the back half on the three places it actually falls over.
The architecture
Two horizontal bands. The top one moves bytes. The bottom one makes decisions. The dashed line between them is the part everyone underestimates.
The data plane
1. Encoder and the ABR ladder
One mezzanine feed comes in from the venue. The encoder fans it out into six to eight renditions — call it 240p at 400 Kbps up to 1080p at 6 Mbps, maybe a 4K rung on top — and chops each into short segments. Two seconds is a common choice.
That segment duration is the single most consequential number in the whole design, and it's a genuine trade. Shorter segments mean lower latency to live, which matters enormously when your viewers can hear their neighbour's TV through the wall. Shorter segments also mean proportionally more requests per viewer per minute, and requests are what the control plane pays for. Going from 6-second to 2-second segments cuts your latency by two thirds and triples your request rate. You do not get one without the other.
2. Packager and origin
The packager wraps each rendition's segments into the delivery formats — CMAF backed by HLS and DASH manifests — and keeps rewriting the manifest as new segments land. Every two seconds, the list of available segments changes.
This is the thing that makes live different from VOD in one sentence: in VOD the manifest is immutable and you can cache it for a year. In live it is the most-requested, fastest-changing object in the system.
3. The origin shield
A single cache tier that sits between every edge PoP and your origin.
Without it, the arithmetic is unkind. If you're delivering from 200 PoPs and each one independently misses on a new segment, your origin takes 200 requests for every segment, per rendition. Multiply by eight renditions and a new segment every two seconds and your origin is fielding 800 requests a second just to serve one stream's worth of unique content.
With a shield, every one of those misses collapses into a single origin fetch. The shield is the cheapest component on the diagram and it's the one standing between you and an origin meltdown.
4. The edge PoPs
Where essentially all of your traffic is actually served. In a healthy live event, north of 99.9% of segment requests never travel past the edge — everyone is watching the same thing at the same time, which is the one genuine gift live gives you. Popular content with perfectly correlated access patterns is the ideal cache workload.
Note the detail in the diagram: the token check happens here too, not in your origin datacenter. I'll come back to why that single decision carries most of the weight.
5. The players
Two million clients, and here's the property that should worry you: they are not independent. They're all locked to the same two-second heartbeat, because they're all following the same manifest describing the same live edge. Independent clients average out. Synchronized clients resonate.
The control plane
6. Identity and entitlement
Answers "is this person allowed to watch this event." Subscription tier, geographic rights, blackout rules, pay-per-view purchase.
The critical design choice is when you answer it. The naive version checks entitlement on every playback request, which puts a database lookup in the path of every viewer. The version that survives issues a short-lived signed token at session start that carries the entitlement claims inside it. The edge then verifies a signature instead of asking a service.
That's the difference between a stateless check at 200 locations and a stateful one in your datacenter.
7. Session and concurrency
"This account is allowed three simultaneous streams" is the requirement that ruins the elegant stateless design, because it is irreducibly stateful — you cannot answer it without shared knowledge of what's currently open.
This is the only service in the whole diagram that holds hard state in the request path, which makes it the one to design most carefully and the first to shed load from. Keep it in a fast keyed store with a short TTL, accept that it will be approximate under stress, and decide in advance which way it fails. Fail-open lets a few people over-share for the duration of an event. Fail-closed locks paying customers out of a broadcast they can never watch again. For a live event, fail-open is almost always the right call — and that should be a deliberate written decision, not an accident of how the timeout happened to be configured.
8. DRM license
Encrypted content needs a license per playback session. The rule that matters: one license per device per event, not per segment. Get this wrong and you've built a request amplifier that fires a few million times in the same ninety seconds.
Persistent licenses scoped to the event duration, issued once and reused for every subsequent segment, turn a sustained flood into a single spike you can plan for.
9. Telemetry and load shedding
QoE beacons come in from players — rebuffer ratio, startup time, which rung of the ladder they settled on. That data is worth having for the post-event report, but the reason it's on the diagram is the orange arrow: it feeds back into the ladder.
When you're in trouble, the most effective lever isn't adding capacity you don't have. It's capping the top rung. Dropping 4K for the duration of the peak can cut egress by a third without a single viewer losing the stream. Build that switch before the event and make sure someone is allowed to pull it without waking an executive.
Where it actually falls over
The manifest stampede
Two million clients refreshing a manifest every two seconds is a million requests per second, on an object that changes constantly.
The fix is unglamorous: cache the manifest at the edge for one second. That sounds pointless — a one-second TTL on a two-second object? — but it's the whole game. It converts a million requests per second into roughly one origin fetch per PoP per second, and it costs you at most one second of additional latency to live.
Pair it with request coalescing so that a thousand simultaneous misses in the same PoP produce one upstream fetch rather than a thousand. Most CDNs do this, but confirm yours is actually enabled for the path rather than assuming.
The auth stampede
This is the one that has genuinely taken services down, because it hides so well in testing. Segment traffic is beautifully cacheable and scales like a dream. Authorization is per-user by definition and caches for nobody.
If every playback start hits a token service in one region, you've built a system where two million people perform a synchronized denial-of-service attack on your own auth tier, while your CDN sits there at 3% utilization wondering what all the fuss is about.
Verify signatures at the edge. A signed token with the entitlement baked in can be checked with a public key at the PoP with no network call at all. Your identity service issues tokens at a rate set by how fast people open the app, which is a gentler curve than how fast they press play, and the verification cost scales with the CDN rather than with your datacenter.
The reconnect storm
The failure mode people forget. Something hiccups — a PoP drops, a transit link flaps — and a few hundred thousand players lose their stream at the same instant. Then every one of them retries. At the same instant. And if that retry path goes through session creation and license acquisition, you get an amplified second wave that is far more expensive than the original one.
Jittered exponential backoff in the player is the fix, and it has to be in the player, which means it has to ship weeks before the event. You cannot deploy a client-side fix during an incident. Of everything in this post, this is the item most likely to be discovered too late.
Some arithmetic worth doing early
Back-of-envelope, for two million concurrent viewers averaging 5 Mbps:
egress 2,000,000 × 5 Mbps = 10 Tbps sustained
segment size 5 Mbps × 2s ≈ 1.25 MB
segment reqs 2,000,000 ÷ 2s = 1M requests/sec
manifest reqs same again = 1M requests/sec
origin fetches 8 renditions ÷ 2s = 4 per second (shielded)
That last line is the one to sit with. The gap between a million requests per second at the edge and four per second at your origin is not a rounding difference — it is the entire architecture, expressed as a ratio. Every decision in the data plane exists to protect it.
What I'd tell someone designing their first one
Separate the two planes explicitly, on paper, before you build. Most live streaming incidents I've seen trace back to a decision that quietly put control-plane work into the data-plane path.
Decide your degradation ladder in advance. When you're ten minutes into a broadcast and the graphs are wrong, you will not design a good mitigation under pressure. Write the steps down while it's calm: cap the ladder, extend the manifest TTL, shed concurrency checks, disable non-critical beacons. In that order.
Load test the synchronized case, not the average one. A test that ramps to two million over an hour tells you almost nothing about two million arriving in ninety seconds. The interesting failures live entirely in the correlation.
Then watch the event with the people who built it. No dashboard teaches you as much as the room does.
Working on something similar, or disagree with where I've put the boundaries? I'd genuinely like to hear it — find me on LinkedIn.
Discussion
Loading…