NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Julian Michael
3 papers
New York University
AI Debate Aids Assessment of Controversial Claims
Quantifying Elicitation of Latent Capabilities in Language Models
Why Do Some Language Models Fake Alignment While Others Don't?