Skip to content
GenAI BR
Trails

How to tell if a model is good

Benchmarks saturate, leaderboards get gamed, long contexts go unused. Six readings to read a launch announcement without being fooled.

6 readings
  1. Reading 1 of 6

    HELM: the proposal to evaluate models across many scenarios and metrics at once, not on a single number

    It is the methodological answer to the single-benchmark problem — and the standard any internal evaluation should copy.

    Evaluationintroductorytechnical, beginnerpaper · Stanford CRFM · TMLR 2023 · 16 Nov 2022
  2. Reading 2 of 6

    A large context window does not mean the context is used: models lose what sits in the middle

    It explains why “dump everything in the context” fails and why RAG and evidence ordering still matter.

    Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023
  3. Reading 3 of 6

    The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data

    If a human-preference ranking can be gamed, “#1 on the Arena” stops being a purchase argument.

    Evaluationtechnicaltechnical, decision-makerpaper · Cohere Labs / Princeton / Stanford / MIT / AI2 · 29 Apr 2025
  4. Reading 4 of 6

    SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks

    It is the number every launch cites; knowing how it is measured is the difference between reading an announcement and understanding it.

    Evaluationtechnicaltechnicalpaper · Princeton / UChicago · ICLR 2024 · 10 Oct 2023
  5. Reading 5 of 6

    Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks

    It is the benchmark that currently discriminates between frontier models on agentic work — and what will appear in the next headlines.

    Evaluationtechnicaltechnicalpaper · Laude Institute / Stanford · 17 Jan 2026
  6. Reading 6 of 6

    GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor

    For builders, it resets the frontier price/performance point; for governance, it is the first flagship whose own report admits reading the reasoning is becoming less reliable.

    Producttechnicaltechnical, decision-makerannouncement · OpenAI · 03 Sept 2026

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.