Mapping the Mind of a Large Language Model
mechanistic-interpretabilityai-safetyfeaturessparse-autoencodersllm
Abstraction: Anthropic extracts millions of interpretable features from Claude 3 Sonnet
Key points:
- Anthropic applied dictionary learning (sparse autoencoders) to the middle layer of Claude 3.0 Sonnet, extracting millions of human-interpretable features — the first detailed look inside a production-grade LLM
- Features are multimodal and multilingual, representing cities, people, elements, code syntax, abstract concepts like "inner conflict," and safety-relevant capabilities (code backdoors, sycophancy, manipulation)
- Feature "distance" reveals conceptual neighborhoods matching human intuitions (e.g., "Golden Gate Bridge" is near Alcatraz, Gavin Newsom, the 1906 earthquake)
- Artificially amplifying a feature causally changes model behavior: activating "Golden Gate Bridge" feature made Claude claim to be the bridge; activating a scam-email feature overcame safety training
- Safety-relevant features found include power-seeking, secrecy, biological weapons, gender bias, and sycophantic praise
- Main limitation: extracting a complete feature set would cost more compute than training the model; circuits (how features are used) remain unknown
Connections: Anthropic · Claude · Mechanistic Interpretability · AI Safety · Dictionary Learning
Source: https://www.anthropic.com/research/mapping-mind-language-model