Llama 3: the 92-page report that showed how a 405B dense model is trained — and what “open” meant in 2024
- What it is
- Meta describes the Llama 3 family (8B, 70B and 405B), dense, with a 128k-token window, trained on about 15 trillion tokens, with details on data, infrastructure, post-training and evaluations that most labs do not publish.
- What was demonstrated
- The 405B was reported as comparable to GPT-4 across several tasks; the family became the base of thousands of fine-tunes and products.
- What was not
- Meta's own license, neither Apache nor MIT — “open weights”, not open source. Comparative evaluations are Meta's own.
- Why it matters
- It is the reference for understanding what changes when a release says Apache 2.0 (Hy4) or MIT (GLM-5.3-Flash) instead of a custom license.
Same theme
Open models→Tencent open-sourced a 770B MoE under Apache 2.0 — and it edges Kimi K3 on code with a quarter of the parameters
GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike
DeepSeek-R1: pure reinforcement learning, with no human reasoning examples, taught the model to verify itself
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.