HELM: the proposal to evaluate models across many scenarios and metrics at once, not on a single number
- What it is
- Stanford's Center for Research on Foundation Models evaluated 30 models across 42 scenarios with 7 metrics — accuracy, calibration, robustness, fairness, bias, toxicity and efficiency — under the same conditions, and published everything.
- What was demonstrated
- Coverage of core scenarios by evaluated models rose from 17.9% to 96.0% — before HELM, most models simply had not been tested on the same things.
- What was not
- From 2022: the scenarios aged and the project became a family of leaderboards. The principle (multi-metric, equal conditions, transparency) is what remains.
- Why it matters
- It is the methodological answer to the single-benchmark problem — and the standard any internal evaluation should copy.
Appears in trails
Same theme
Evaluation→On real scientific software, the best coding agent solves fewer than half the bugs
SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
A large context window does not mean the context is used: models lose what sits in the middle
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.