Skip to content
GenAI BR
Archive · Theme7 entries in this theme

Evaluation

Benchmarks, contamination, leaderboards and the gap between the score and real use.

  1. On real scientific software, the best coding agent solves fewer than half the bugs

    It quantifies the gap between saturated leaderboards and real domain-heavy engineering — exactly where engineers will be asked to deploy agents next.

    Evaluationtechnicaltechnicalpaper · Universidade Fudan · 01 Sept 2026
  2. SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks

    It is the number every launch cites; knowing how it is measured is the difference between reading an announcement and understanding it.

    Evaluationtechnicaltechnicalpaper · Princeton / UChicago · ICLR 2024 · 10 Oct 2023
  3. Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks

    It is the benchmark that currently discriminates between frontier models on agentic work — and what will appear in the next headlines.

    Evaluationtechnicaltechnicalpaper · Laude Institute / Stanford · 17 Jan 2026
  4. A large context window does not mean the context is used: models lose what sits in the middle

    It explains why “dump everything in the context” fails and why RAG and evidence ordering still matter.

    Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023
  5. The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data

    If a human-preference ranking can be gamed, “#1 on the Arena” stops being a purchase argument.

    Evaluationtechnicaltechnical, decision-makerpaper · Cohere Labs / Princeton / Stanford / MIT / AI2 · 29 Apr 2025
  6. HELM: the proposal to evaluate models across many scenarios and metrics at once, not on a single number

    It is the methodological answer to the single-benchmark problem — and the standard any internal evaluation should copy.

    Evaluationintroductorytechnical, beginnerpaper · Stanford CRFM · TMLR 2023 · 16 Nov 2022
  7. DeepSeek-R1: pure reinforcement learning, with no human reasoning examples, taught the model to verify itself

    It defined the open “reasoning model” recipe that China has been iterating on ever since.

    Open modelstechnicaltechnicalpaper · DeepSeek · Nature 645 (2025) · 22 Jan 2025