Mapping the Mind of a Large Language Model

mechanistic-interpretabilityai-safetyfeaturessparse-autoencodersllm

Abstraction: Anthropic extracts millions of interpretable features from Claude 3 Sonnet

Key points:

Connections: Anthropic · Claude · Mechanistic Interpretability · AI Safety · Dictionary Learning

Source: https://www.anthropic.com/research/mapping-mind-language-model