Skip to content
GenAI BR
Infra and costIn the archive since 06 Sept 2026

PagedAttention: treating the KV cache like virtual memory doubled inference throughput — and gave birth to vLLM

Infra and costtechnicaltechnicalpaper · UC Berkeley · SOSP 2023 · 12 Sept 2023
Download card
What it is
Kwon, Li and colleagues (Berkeley) observed that the attention key-value cache wasted memory through fragmentation, and applied the operating-system idea of paging: non-contiguous, shareable blocks across requests. The result is vLLM, today one of the most used inference servers.
What was demonstrated
2–4× higher throughput than FasterTransformer and Orca at the same latency, with larger gains for longer sequences and bigger models.
What was not
2023 numbers; the competitive landscape (SGLang, TensorRT-LLM) has changed. The principle, not the comparison, is what holds.
Why it matters
Cost per token is largely memory cost; this is the paper that explains why.
  1. GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens

    Infra and costtechnicaltechnical, decision-makerannouncement · Z.ai · 26 Aug 2026
  2. The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October

    Regulation and Brazilintroductorytechnical, decision-maker, beginnerannouncement · Governo Federal · 21 Aug 2026
  3. A large context window does not mean the context is used: models lose what sits in the middle

    Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023
  4. DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike

    Infra and costtechnicaltechnical, decision-makerpaper · DeepSeek · 27 Dec 2024

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.