Mechanistic Interpretability

concepts · 2 notes linked

Related: GPT-2 · Openai · GPT-3 · Transformers · Superposition · Anthropic · Claude · AI Safety

Notes