NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
harmful content
3 papers
One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention Head
Safety Depth in Large Language Models: A Markov Chain Perspective
Safety Pretraining: Toward the Next Generation of Safe AI