DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike
- What it is
- The technical report describes a 671-billion-parameter MoE (37B active) with Multi-head Latent Attention, auxiliary-loss-free load balancing and multi-token prediction, trained on 14.8 trillion tokens on H800 GPUs — the export-restricted version for China.
- What was demonstrated
- 2.788 million GPU-hours of training, with no unrecoverable loss spikes or rollbacks; performance comparable to closed models of the time at a fraction of the estimated cost.
- What was not
- The GPU-hour figure covers the final run, not prior experiments or infrastructure. Training data is not described in detail.
- Why it matters
- It was the report that made the market rethink what a frontier model costs to train — and the origin of several techniques 2026 open models still use.
Same theme
Infra and cost→GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October
A large context window does not mean the context is used: models lose what sits in the middle
PagedAttention: treating the KV cache like virtual memory doubled inference throughput — and gave birth to vLLM
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.