Millions of interpretable features extracted from a production model — including one that makes it think it is the Golden Gate Bridge
- What it is
- Anthropic's interpretability team trained sparse autoencoders with 1, 4 and 34 million features on Claude 3 Sonnet's activations, and showed that many correspond to readable concepts — places, people, insecure code, sycophancy — and that artificially activating them changes the model's behavior.
- What was demonstrated
- Multilingual and multimodal features; safety-relevant features (deception, power-seeking, bias); the public “Golden Gate Claude” demo, where a single amplified feature dominated responses.
- What was not
- Published by the lab about its own model, without external review. Sparse autoencoders capture part of the activations; what is left out is an open question.
- Why it matters
- It was the moment mechanistic interpretability moved from toy models to a model people actually use.
Same theme
Interpretability→Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.