sparse autoencoders
A type of neural network that is trained to encode input data into a lower-dimensional space, with the constraint that only a small number of neurons are activated at any time. This leads to a more efficient representation that captures the essential features of the input data.
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
- Among Us: A Sandbox for Measuring and Detecting Agentic Deception
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- Dense SAE Latents Are Features, Not Bugs
- Emergence and Evolution of Interpretable Concepts in Diffusion Models
- From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Interpreting vision transformers via residual replacement model
- LLM Interpretability with Identifiable Temporal-Instantaneous Representation
- LLM Layers Immediately Correct Each Other
- Large Language Models Think Too Fast To Explore Effectively
- Measuring and Guiding Monosemanticity
- One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
- Revising and Falsifying Sparse Autoencoder Feature Explanations
- SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
- Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations
- Transferring Linear Features Across Language Models With Model Stitching