Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
- What it is
- Merrill, Shaw, Carlini and over sixty co-authors (Laude Institute, Stanford and others) published in January 2026 a set of realistic command-line tasks — environment setup, debugging, compiling, data processing — with test-based verification. The benchmark is “continuous”: the current version on Harbor Hub is 4.0.
- What was demonstrated
- At publication (2.0), the best agents scored under 65%. In September 2026, vendors quote 4.0: GPT-6 Astra 57.9%, Claude Fable 5.1 55.8%.
- What was not
- Versions change content; comparing numbers across versions is a common error. Vendor numbers are self-reported.
- Why it matters
- It is the benchmark that currently discriminates between frontier models on agentic work — and what will appear in the next headlines.
Appears in trails
Same theme
Evaluation→On real scientific software, the best coding agent solves fewer than half the bugs
SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
A large context window does not mean the context is used: models lose what sits in the middle
The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.