Skip to content
GenAI BR
EvaluationFeatured in Cinco no. 1 · In the archive since 08 Sept 2026

On real scientific software, the best coding agent solves fewer than half the bugs

Evaluationtechnicaltechnicalpaper · Universidade Fudan · 01 Sept 2026
Download card
What it is
SWE-bench Science, from Fudan University's OpenMOSS group, collects 119 real issues from 98 repositories across 20 scientific domains — imaging, spectroscopy, simulation. Failures involve units, coordinate systems, numerical invariants and physical assumptions, not plain logic. The authors propose a four-mechanism failure taxonomy and test whether injecting scientific guidance helps.
What was demonstrated
Leaderboard on 1 September: Claude Opus 5 (max effort) 47.9%; DeepSeek-V4-Pro 42.0%; GPT-5.6 Sol 40.3%; Kimi K3 35.3%; GLM-5.2 31.9%. The same models score above 95% on SWE-bench Verified. Well-grounded guidance helps; poorly calibrated guidance causes anchoring and hurts.
What was not
Small sample (119 tasks): differences of a few points between models are within noise. Academic group with no vendor ties, but run cost limits replication. Preprint v2, not yet peer-reviewed. Not to be confused with “Terminal-Bench-Science 0.1”, which appears in this week's launch decks.
Why it matters
It quantifies the gap between saturated leaderboards and real domain-heavy engineering — exactly where engineers will be asked to deploy agents next.
Appears in trails

Same theme

Evaluation
  1. SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks

    Evaluationtechnicaltechnicalpaper · Princeton / UChicago · ICLR 2024 · 10 Oct 2023
  2. Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks

    Evaluationtechnicaltechnicalpaper · Laude Institute / Stanford · 17 Jan 2026
  3. A large context window does not mean the context is used: models lose what sits in the middle

    Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023
  4. The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data

    Evaluationtechnicaltechnical, decision-makerpaper · Cohere Labs / Princeton / Stanford / MIT / AI2 · 29 Apr 2025

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.