activation patterns
- Correcting misinterpretations of additive models
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- Optimal Rates for Generalization of Gradient Descent for Deep ReLU Classification
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons