DeepSeek-R1: pure reinforcement learning, with no human reasoning examples, taught the model to verify itself
- What it is
- R1-Zero was trained only with reinforcement learning on verifiable rewards (correct answers in math and code), without supervised reasoning data; R1 adds a cold-start data stage and further training. The work was published in Nature in 2025, and the weights are open.
- What was demonstrated
- Self-reflection and verification behaviors emerged from reinforcement; versions distilled to small models (1.5B to 70B) transferred part of the capability.
- What was not
- R1-Zero mixed languages and was barely readable; the recipe requires verifiable rewards, which limits the method to domains with a right answer. Inference cost of long reasoning is high.
- Why it matters
- It defined the open “reasoning model” recipe that China has been iterating on ever since.
Same theme
Open models→Tencent open-sourced a 770B MoE under Apache 2.0 — and it edges Kimi K3 on code with a quarter of the parameters
GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike
Llama 3: the 92-page report that showed how a 405B dense model is trained — and what “open” meant in 2024
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.