NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Fabien Roger
3 papers
Anthropic
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
Quantifying Elicitation of Latent Capabilities in Language Models
Why Do Some Language Models Fake Alignment While Others Don't?