Skip to content
GenAI BR
Archive · Theme5 entries in this theme

Open models

Open weights, licenses, technical reports and what you can run yourself.

  1. Tencent open-sourced a 770B MoE under Apache 2.0 — and it edges Kimi K3 on code with a quarter of the parameters

    It is the largest Apache-2.0 open-weight release of the period, and it confirms the pattern: Chinese labs shipping frontier-adjacent agentic coders at a fraction of proprietary cost.

    Open modelstechnicaltechnical, decision-makerannouncement · Tencent · 28 Aug 2026
  2. GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens

    It is the clearest datapoint yet that a competitive model can be trained and served at scale without NVIDIA hardware — and it prices Opus-class agentic work at cents.

    Infra and costtechnicaltechnical, decision-makerannouncement · Z.ai · 26 Aug 2026
  3. DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike

    It was the report that made the market rethink what a frontier model costs to train — and the origin of several techniques 2026 open models still use.

    Infra and costtechnicaltechnical, decision-makerpaper · DeepSeek · 27 Dec 2024
  4. DeepSeek-R1: pure reinforcement learning, with no human reasoning examples, taught the model to verify itself

    It defined the open “reasoning model” recipe that China has been iterating on ever since.

    Open modelstechnicaltechnicalpaper · DeepSeek · Nature 645 (2025) · 22 Jan 2025
  5. Llama 3: the 92-page report that showed how a 405B dense model is trained — and what “open” meant in 2024

    It is the reference for understanding what changes when a release says Apache 2.0 (Hy4) or MIT (GLM-5.3-Flash) instead of a custom license.

    Open modelsintroductorytechnical, beginnerpaper · Meta · 31 Jul 2024