Skip to content
GenAI BR
InterpretabilityIn the archive since 06 Sept 2026

Attribution graphs show the model planning rhymes ahead — and chains of thought that do not match the real computation

Interpretabilityadvancedtechnicalpaper · Anthropic · 27 Mar 2025
Download card
What it is
“On the Biology of a Large Language Model”, with its methods companion “Circuit Tracing”, applies attribution graphs to Claude 3.5 Haiku to trace computational circuits on concrete tasks: multi-hop reasoning, rhyme planning in poetry, shared multilingual circuits, arithmetic.
What was demonstrated
Evidence of forward planning (the model picks the rhyme word before writing the line) and cases where the chain-of-thought explanation is unfaithful to what the model actually computed.
What was not
Case studies, not statistics; published by the lab about its own model. The authors list the method's limitations at length.
Why it matters
Chain-of-thought unfaithfulness is the technical reason why “reading the reasoning” is no safety guarantee — which the GPT-6 Astra system card just admitted in practice.
  1. GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor

    Producttechnicaltechnical, decision-makerannouncement · OpenAI · 03 Sept 2026
  2. Millions of interpretable features extracted from a production model — including one that makes it think it is the Golden Gate Bridge

    Interpretabilityadvancedtechnicalpaper · Anthropic · 21 May 2024

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.