Superposition
concepts · 1 notes linked
Related: GPT-2 · Openai · GPT-3 · Mechanistic Interpretability · Transformers
Notes
- The Singular Value Decompositions of Transformer Weight Matrices — SVD of GPT-2 weight matrices reveals interpretable semantic directions