Skip to content
GenAI BR
EvaluationIn the archive since 06 Sept 2026

A large context window does not mean the context is used: models lose what sits in the middle

Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023
Download card
What it is
Liu and colleagues (Stanford, Berkeley, Samaya) measured how models use long documents by moving the relevant information around inside the context. Performance draws a U: high when the information is at the start or end, and it drops when it sits in the middle.
What was demonstrated
Significant accuracy drop on multi-document QA and key-value retrieval when the relevant item is in the middle, even in models sold as “long-context”.
What was not
Published in 2023; current models reduced the effect but did not eliminate it. Measure in your own case before trusting the advertised window.
Why it matters
It explains why “dump everything in the context” fails and why RAG and evidence ordering still matter.
Appears in trails

Same theme

Evaluation
  1. On real scientific software, the best coding agent solves fewer than half the bugs

    Evaluationtechnicaltechnicalpaper · Universidade Fudan · 01 Sept 2026
  2. SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks

    Evaluationtechnicaltechnicalpaper · Princeton / UChicago · ICLR 2024 · 10 Oct 2023
  3. Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks

    Evaluationtechnicaltechnicalpaper · Laude Institute / Stanford · 17 Jan 2026
  4. The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data

    Evaluationtechnicaltechnical, decision-makerpaper · Cohere Labs / Princeton / Stanford / MIT / AI2 · 29 Apr 2025

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.