How to tell if a model is good
Benchmarks saturate, leaderboards get gamed, long contexts go unused. Six readings to read a launch announcement without being fooled.
HELM: the proposal to evaluate models across many scenarios and metrics at once, not on a single number
It is the methodological answer to the single-benchmark problem — and the standard any internal evaluation should copy.
A large context window does not mean the context is used: models lose what sits in the middle
It explains why “dump everything in the context” fails and why RAG and evidence ordering still matter.
The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
If a human-preference ranking can be gamed, “#1 on the Arena” stops being a purchase argument.
SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
It is the number every launch cites; knowing how it is measured is the difference between reading an announcement and understanding it.
Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
It is the benchmark that currently discriminates between frontier models on agentic work — and what will appear in the next headlines.
GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor
For builders, it resets the frontier price/performance point; for governance, it is the first flagship whose own report admits reading the reasoning is becoming less reliable.
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.