Archive · Theme
Interpretability
What happens inside the model: features, circuits, chain-of-thought faithfulness.
GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor
For builders, it resets the frontier price/performance point; for governance, it is the first flagship whose own report admits reading the reasoning is becoming less reliable.
Millions of interpretable features extracted from a production model — including one that makes it think it is the Golden Gate Bridge
It was the moment mechanistic interpretability moved from toy models to a model people actually use.
Attribution graphs show the model planning rhymes ahead — and chains of thought that do not match the real computation
Chain-of-thought unfaithfulness is the technical reason why “reading the reasoning” is no safety guarantee — which the GPT-6 Astra system card just admitted in practice.