A large context window does not mean the context is used: models lose what sits in the middle
- What it is
- Liu and colleagues (Stanford, Berkeley, Samaya) measured how models use long documents by moving the relevant information around inside the context. Performance draws a U: high when the information is at the start or end, and it drops when it sits in the middle.
- What was demonstrated
- Significant accuracy drop on multi-document QA and key-value retrieval when the relevant item is in the middle, even in models sold as “long-context”.
- What was not
- Published in 2023; current models reduced the effect but did not eliminate it. Measure in your own case before trusting the advertised window.
- Why it matters
- It explains why “dump everything in the context” fails and why RAG and evidence ordering still matter.
Appears in trails
Same theme
Evaluation→On real scientific software, the best coding agent solves fewer than half the bugs
SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.