Skip to content
GenAI BR
InterpretabilityIn the archive since 06 Sept 2026

Millions of interpretable features extracted from a production model — including one that makes it think it is the Golden Gate Bridge

Interpretabilityadvancedtechnicalpaper · Anthropic · 21 May 2024
Download card
What it is
Anthropic's interpretability team trained sparse autoencoders with 1, 4 and 34 million features on Claude 3 Sonnet's activations, and showed that many correspond to readable concepts — places, people, insecure code, sycophancy — and that artificially activating them changes the model's behavior.
What was demonstrated
Multilingual and multimodal features; safety-relevant features (deception, power-seeking, bias); the public “Golden Gate Claude” demo, where a single amplified feature dominated responses.
What was not
Published by the lab about its own model, without external review. Sparse autoencoders capture part of the activations; what is left out is an open question.
Why it matters
It was the moment mechanistic interpretability moved from toy models to a model people actually use.
  1. GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor

    Producttechnicaltechnical, decision-makerannouncement · OpenAI · 03 Sept 2026
  2. Attribution graphs show the model planning rhymes ahead — and chains of thought that do not match the real computation

    Interpretabilityadvancedtechnicalpaper · Anthropic · 27 Mar 2025

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.