NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
model activations
3 papers
Detecting High-Stakes Interactions with Activation Probes
On Reasoning Strength Planning in Large Reasoning Models
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry