On real scientific software, the best coding agent solves fewer than half the bugs
- What it is
- SWE-bench Science, from Fudan University's OpenMOSS group, collects 119 real issues from 98 repositories across 20 scientific domains — imaging, spectroscopy, simulation. Failures involve units, coordinate systems, numerical invariants and physical assumptions, not plain logic. The authors propose a four-mechanism failure taxonomy and test whether injecting scientific guidance helps.
- What was demonstrated
- Leaderboard on 1 September: Claude Opus 5 (max effort) 47.9%; DeepSeek-V4-Pro 42.0%; GPT-5.6 Sol 40.3%; Kimi K3 35.3%; GLM-5.2 31.9%. The same models score above 95% on SWE-bench Verified. Well-grounded guidance helps; poorly calibrated guidance causes anchoring and hurts.
- What was not
- Small sample (119 tasks): differences of a few points between models are within noise. Academic group with no vendor ties, but run cost limits replication. Preprint v2, not yet peer-reviewed. Not to be confused with “Terminal-Bench-Science 0.1”, which appears in this week's launch decks.
- Why it matters
- It quantifies the gap between saturated leaderboards and real domain-heavy engineering — exactly where engineers will be asked to deploy agents next.
Appears in trails
Same theme
Evaluation→SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
A large context window does not mean the context is used: models lose what sits in the middle
The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.