Attribution graphs show the model planning rhymes ahead — and chains of thought that do not match the real computation
- What it is
- “On the Biology of a Large Language Model”, with its methods companion “Circuit Tracing”, applies attribution graphs to Claude 3.5 Haiku to trace computational circuits on concrete tasks: multi-hop reasoning, rhyme planning in poetry, shared multilingual circuits, arithmetic.
- What was demonstrated
- Evidence of forward planning (the model picks the rhyme word before writing the line) and cases where the chain-of-thought explanation is unfaithful to what the model actually computed.
- What was not
- Case studies, not statistics; published by the lab about its own model. The authors list the method's limitations at length.
- Why it matters
- Chain-of-thought unfaithfulness is the technical reason why “reading the reasoning” is no safety guarantee — which the GPT-6 Astra system card just admitted in practice.
Same theme
Interpretability→Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.