mechanistic interpretability
Mechanistic interpretability focuses on understanding the decision-making processes within AI models at a granular level, elucidating how various components contribute to outputs.
- A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning
- Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
- Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
- DePass: Unified Feature Attributing by Simple Decomposed Forward Pass
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
- Emergence and Evolution of Interpretable Concepts in Diffusion Models
- Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
- LLM Interpretability with Identifiable Temporal-Instantaneous Representation
- Measuring and Guiding Monosemanticity
- Mitigating Overthinking in Large Reasoning Models via Manifold Steering
- Prompting as Scientific Inquiry
- Proxy-SPEX: Sample-Efficient Interpretability via Sparse Feature Interactions in LLMs
- Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates
- Revising and Falsifying Sparse Autoencoder Feature Explanations
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons