NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Ekdeep S Lubana
4 papers
Goodfire AI
Detecting High-Stakes Interactions with Activation Probes
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
In-Context Learning Strategies Emerge Rationally
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry