PagedAttention: treating the KV cache like virtual memory doubled inference throughput — and gave birth to vLLM
- What it is
- Kwon, Li and colleagues (Berkeley) observed that the attention key-value cache wasted memory through fragmentation, and applied the operating-system idea of paging: non-contiguous, shareable blocks across requests. The result is vLLM, today one of the most used inference servers.
- What was demonstrated
- 2–4× higher throughput than FasterTransformer and Orca at the same latency, with larger gains for longer sequences and bigger models.
- What was not
- 2023 numbers; the competitive landscape (SGLang, TensorRT-LLM) has changed. The principle, not the comparison, is what holds.
- Why it matters
- Cost per token is largely memory cost; this is the paper that explains why.
Same theme
Infra and cost→GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October
A large context window does not mean the context is used: models lose what sits in the middle
DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.