Dictionary Learning
concepts · 1 notes linked
Related: Anthropic · Claude · Mechanistic Interpretability · AI Safety
Notes
- Mapping the Mind of a Large Language Model — Anthropic extracts millions of interpretable features from Claude 3 Sonnet
concepts · 1 notes linked
Related: Anthropic · Claude · Mechanistic Interpretability · AI Safety