Skip to content
GenAI BR
EvaluationIn the archive since 06 Sept 2026

SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks

Evaluationtechnicaltechnicalpaper · Princeton / UChicago · ICLR 2024 · 10 Oct 2023
Download card
What it is
Jimenez and colleagues (Princeton and Chicago) assembled 2,294 real issues from 12 Python repositories, with tests that define whether the fix works. In 2024 OpenAI published the “Verified” subset, 500 human-validated tasks, which became the reference number in launches.
What was demonstrated
At publication the best model (Claude 2) solved 1.96%. In September 2026, independent Vals.ai runs put Claude Opus 5, GPT-5.6 Sol and GPT-5.6 Terra at 97.0% on Verified; self-reported aggregators disagree on the ordering.
What was not
Saturated: with everyone near 97%, the benchmark no longer separates models. Results depend on the harness (which agent, how many attempts), and the official leaderboard mixes self-report and verification. Python only.
Why it matters
It is the number every launch cites; knowing how it is measured is the difference between reading an announcement and understanding it.
Appears in trails

Same theme

Evaluation
  1. On real scientific software, the best coding agent solves fewer than half the bugs

    Evaluationtechnicaltechnicalpaper · Universidade Fudan · 01 Sept 2026
  2. Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks

    Evaluationtechnicaltechnicalpaper · Laude Institute / Stanford · 17 Jan 2026
  3. A large context window does not mean the context is used: models lose what sits in the middle

    Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023
  4. The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data

    Evaluationtechnicaltechnical, decision-makerpaper · Cohere Labs / Princeton / Stanford / MIT / AI2 · 29 Apr 2025

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.