The ISP Model

An ISP sells 100 households “unlimited 10 Gbps” on a shared 10 Gbps pipe. They bet that at any given moment, only a fraction of those households are actively using bandwidth. The math holds because: - Most browsing uses < 5 Mbps - Peak hour lasts 2-3 hours - Even “heavy” use is bursty

The LLM Model

Now look at a GPT-5.6 inference API. They provision GPUs that can handle X tokens per second. They sell “unlimited tokens” or generous rate limits to Y customers where Y × typical_usage ≫ X.

The bet: - Most users send short queries - Peak hours are predictable (work hours in their region) - Even “heavy” users average far below their peak burst - Real-time inference is bursty (you think for 10 seconds, then send a prompt in 2 seconds)

Key Parallel

Concept ISP World LLM World
The pipe GPON port / submarine cable GPU cluster / inference server
The product Unlimited 10 Gbps Unlimited tokens
The bet Most users are idle Most users are thinking
Concept | ISP World | LLM World |
Concept | ISP World | LLM World |
Concept | ISP World | LLM World |
Concept ISP World LLM World
Concept ISP World LLM World
Concept ISP World LLM World
Concept ISP World LLM World
Peak time 8 PM weekdays US/EU business hours
Contention 24:1 (Simba) Unknown, but similar
Hardware cap Wi-Fi / Ethernet Token context window
What you actually get 10 Gbps / 24 = 400 Mbps X tokens / Y users

The Hardware Cap — You Can’t Consume It

Just as Wi-Fi caps your real bandwidth at ~270 Mbps regardless of plan, your reading speed caps your token consumption. Even with the fastest LLM:

Your consumption rate is far below what the API can deliver. The provider knows this.

The Economic Incentive Is the Same

ISP: - Sell big plans people don’t need - Profit from users who never saturate - Raise prices only when contention breaks

LLM Provider: - Sell “unlimited” API tiers - Profit from users who hit rate limits rarely - Introduce tiered pricing when usage spikes

Both are selling headroom based on the statistical reality that most customers don’t use what they pay for.

The Difference

There is one difference: LLM inference can’t be as aggressively oversubscribed as bandwidth because inference is compute-bound not pipe-bound.

A GPON port can handle 24 homes because idle customers use near-zero bandwidth. LLM inference GPUs idle at near-zero utilization too — but the burst when a customer sends a prompt is much more intense relative to capacity. A single ChatGPT query can use hundreds of GPU-seconds.

So LLM providers are more cautious with oversubscription. But the principle is identical: they’re selling you access to a shared pool, not dedicated capacity.

What This Means

When you see “unlimited tokens” or “unlimited broadband,” mentally translate it to: “as many as we can give you before you notice the contention.”

For broadband, you don’t notice because your Wi-Fi caps you before the contention does.

For LLMs, you don’t notice because you read slower than the model generates.

In both cases, the bottleneck is you — not the provider.

Related: Oversubscription and “Unlimited” Marketing — The Hidden Math