Skip to content
GenAI BR
Archive · Theme5 entries in this theme

Infra and cost

Inference, GPUs, price per token, and what changes when cost drops tenfold.

  1. GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens

    It is the clearest datapoint yet that a competitive model can be trained and served at scale without NVIDIA hardware — and it prices Opus-class agentic work at cents.

    Infra and costtechnicaltechnical, decision-makerannouncement · Z.ai · 26 Aug 2026
  2. The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October

    It is the first concrete, dated and funded step toward sovereign AI compute in Brazil — and the tender terms will shape which stacks Brazilian engineers get to use.

    Regulation and Brazilintroductorytechnical, decision-maker, beginnerannouncement · Governo Federal · 21 Aug 2026
  3. A large context window does not mean the context is used: models lose what sits in the middle

    It explains why “dump everything in the context” fails and why RAG and evidence ordering still matter.

    Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023
  4. DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike

    It was the report that made the market rethink what a frontier model costs to train — and the origin of several techniques 2026 open models still use.

    Infra and costtechnicaltechnical, decision-makerpaper · DeepSeek · 27 Dec 2024
  5. PagedAttention: treating the KV cache like virtual memory doubled inference throughput — and gave birth to vLLM

    Cost per token is largely memory cost; this is the paper that explains why.

    Infra and costtechnicaltechnicalpaper · UC Berkeley · SOSP 2023 · 12 Sept 2023