Evaluation
Benchmarks, contamination, leaderboards and the gap between the score and real use.
On real scientific software, the best coding agent solves fewer than half the bugs
It quantifies the gap between saturated leaderboards and real domain-heavy engineering — exactly where engineers will be asked to deploy agents next.
SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
It is the number every launch cites; knowing how it is measured is the difference between reading an announcement and understanding it.
Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
It is the benchmark that currently discriminates between frontier models on agentic work — and what will appear in the next headlines.
A large context window does not mean the context is used: models lose what sits in the middle
It explains why “dump everything in the context” fails and why RAG and evidence ordering still matter.
The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
If a human-preference ranking can be gamed, “#1 on the Arena” stops being a purchase argument.
HELM: the proposal to evaluate models across many scenarios and metrics at once, not on a single number
It is the methodological answer to the single-benchmark problem — and the standard any internal evaluation should copy.
DeepSeek-R1: pure reinforcement learning, with no human reasoning examples, taught the model to verify itself
It defined the open “reasoning model” recipe that China has been iterating on ever since.