Open models
Open weights, licenses, technical reports and what you can run yourself.
Tencent open-sourced a 770B MoE under Apache 2.0 — and it edges Kimi K3 on code with a quarter of the parameters
It is the largest Apache-2.0 open-weight release of the period, and it confirms the pattern: Chinese labs shipping frontier-adjacent agentic coders at a fraction of proprietary cost.
GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
It is the clearest datapoint yet that a competitive model can be trained and served at scale without NVIDIA hardware — and it prices Opus-class agentic work at cents.
DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike
It was the report that made the market rethink what a frontier model costs to train — and the origin of several techniques 2026 open models still use.
DeepSeek-R1: pure reinforcement learning, with no human reasoning examples, taught the model to verify itself
It defined the open “reasoning model” recipe that China has been iterating on ever since.
Llama 3: the 92-page report that showed how a 405B dense model is trained — and what “open” meant in 2024
It is the reference for understanding what changes when a release says Apache 2.0 (Hy4) or MIT (GLM-5.3-Flash) instead of a custom license.