Mechanistic Interpretability
concepts · 2 notes linked
Related: GPT-2 · Openai · GPT-3 · Transformers · Superposition · Anthropic · Claude · AI Safety
Notes
- Mapping the Mind of a Large Language Model — Anthropic extracts millions of interpretable features from Claude 3 Sonnet
- The Singular Value Decompositions of Transformer Weight Matrices — SVD of GPT-2 weight matrices reveals interpretable semantic directions