Skip to content
GenAI BR
EvaluationIn the archive since 06 Sept 2026

HELM: the proposal to evaluate models across many scenarios and metrics at once, not on a single number

Evaluationintroductorytechnical, beginnerpaper · Stanford CRFM · TMLR 2023 · 16 Nov 2022
Download card
What it is
Stanford's Center for Research on Foundation Models evaluated 30 models across 42 scenarios with 7 metrics — accuracy, calibration, robustness, fairness, bias, toxicity and efficiency — under the same conditions, and published everything.
What was demonstrated
Coverage of core scenarios by evaluated models rose from 17.9% to 96.0% — before HELM, most models simply had not been tested on the same things.
What was not
From 2022: the scenarios aged and the project became a family of leaderboards. The principle (multi-metric, equal conditions, transparency) is what remains.
Why it matters
It is the methodological answer to the single-benchmark problem — and the standard any internal evaluation should copy.
Appears in trails

Same theme

Evaluation
  1. On real scientific software, the best coding agent solves fewer than half the bugs

    Evaluationtechnicaltechnicalpaper · Universidade Fudan · 01 Sept 2026
  2. SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks

    Evaluationtechnicaltechnicalpaper · Princeton / UChicago · ICLR 2024 · 10 Oct 2023
  3. Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks

    Evaluationtechnicaltechnicalpaper · Laude Institute / Stanford · 17 Jan 2026
  4. A large context window does not mean the context is used: models lose what sits in the middle

    Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.