The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
- What it is
- Singh, Hooker and co-authors (Cohere Labs, Princeton, Stanford, MIT, AI2) analyzed Chatbot Arena and documented practices that distort the ranking: private testing of multiple variants before launch, selective withdrawal of results, and unequal access to evaluation data.
- What was demonstrated
- Meta tested 27 private variants before Llama 4; Google and OpenAI received ~19% and ~20% of all Arena data, versus ~30% for 83 open models combined; extra Arena data yielded up to +112% gains on the Arena's own distribution.
- What was not
- Several authors work at Cohere, which competes on the Arena. The Arena organization contested part of the methodology and changed policies after publication.
- Why it matters
- If a human-preference ranking can be gamed, “#1 on the Arena” stops being a purchase argument.
Appears in trails
Same theme
Evaluation→On real scientific software, the best coding agent solves fewer than half the bugs
SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
A large context window does not mean the context is used: models lose what sits in the middle
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.