Unlimited Bandwidth and Unlimited Tokens — They're the Same Lie
The ISP Model
An ISP sells 100 households “unlimited 10 Gbps” on a shared 10 Gbps pipe. They bet that at any given moment, only a fraction of those households are actively using bandwidth. The math holds because: - Most browsing uses < 5 Mbps - Peak hour lasts 2-3 hours - Even “heavy” use is bursty
The LLM Model
Now look at a GPT-5.6 inference API. They provision GPUs that can handle X tokens per second. They sell “unlimited tokens” or generous rate limits to Y customers where Y × typical_usage ≫ X.
The bet: - Most users send short queries - Peak hours are predictable (work hours in their region) - Even “heavy” users average far below their peak burst - Real-time inference is bursty (you think for 10 seconds, then send a prompt in 2 seconds)
Key Parallel
| Concept | ISP World | LLM World |
|---|---|---|
| The pipe | GPON port / submarine cable | GPU cluster / inference server |
| The product | Unlimited 10 Gbps | Unlimited tokens |
| The bet | Most users are idle | Most users are thinking |
| Concept | ISP World | LLM World |
|---|---|---|
| — | — | — |
| Concept | ISP World | LLM World |
| — | — | — |
| — | — | — |
| Concept | ISP World | LLM World |
| Concept | ISP World | LLM World |
|---|---|---|
| — | — | — |
| — | — | — |
| Peak time | 8 PM weekdays | US/EU business hours |
| Contention | 24:1 (Simba) | Unknown, but similar |
| Hardware cap | Wi-Fi / Ethernet | Token context window |
| What you actually get | 10 Gbps / 24 = 400 Mbps | X tokens / Y users |
The Hardware Cap — You Can’t Consume It
Just as Wi-Fi caps your real bandwidth at ~270 Mbps regardless of plan, your reading speed caps your token consumption. Even with the fastest LLM:
- A human reads at ~200-300 words per minute
- A good LLM generates ~50-100 tokens per second
- A 1,000-token response takes 10-20 seconds to generate
- You then spend 30-60 seconds reading it
Your consumption rate is far below what the API can deliver. The provider knows this.
The Economic Incentive Is the Same
ISP: - Sell big plans people don’t need - Profit from users who never saturate - Raise prices only when contention breaks
LLM Provider: - Sell “unlimited” API tiers - Profit from users who hit rate limits rarely - Introduce tiered pricing when usage spikes
Both are selling headroom based on the statistical reality that most customers don’t use what they pay for.
The Difference
There is one difference: LLM inference can’t be as aggressively oversubscribed as bandwidth because inference is compute-bound not pipe-bound.
A GPON port can handle 24 homes because idle customers use near-zero bandwidth. LLM inference GPUs idle at near-zero utilization too — but the burst when a customer sends a prompt is much more intense relative to capacity. A single ChatGPT query can use hundreds of GPU-seconds.
So LLM providers are more cautious with oversubscription. But the principle is identical: they’re selling you access to a shared pool, not dedicated capacity.
What This Means
When you see “unlimited tokens” or “unlimited broadband,” mentally translate it to: “as many as we can give you before you notice the contention.”
For broadband, you don’t notice because your Wi-Fi caps you before the contention does.
For LLMs, you don’t notice because you read slower than the model generates.
In both cases, the bottleneck is you — not the provider.
Related: Oversubscription and “Unlimited” Marketing — The Hidden Math