SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
- What it is
- Jimenez and colleagues (Princeton and Chicago) assembled 2,294 real issues from 12 Python repositories, with tests that define whether the fix works. In 2024 OpenAI published the “Verified” subset, 500 human-validated tasks, which became the reference number in launches.
- What was demonstrated
- At publication the best model (Claude 2) solved 1.96%. In September 2026, independent Vals.ai runs put Claude Opus 5, GPT-5.6 Sol and GPT-5.6 Terra at 97.0% on Verified; self-reported aggregators disagree on the ordering.
- What was not
- Saturated: with everyone near 97%, the benchmark no longer separates models. Results depend on the harness (which agent, how many attempts), and the official leaderboard mixes self-report and verification. Python only.
- Why it matters
- It is the number every launch cites; knowing how it is measured is the difference between reading an announcement and understanding it.
Appears in trails
Same theme
Evaluation→On real scientific software, the best coding agent solves fewer than half the bugs
Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
A large context window does not mean the context is used: models lose what sits in the middle
The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.