NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
model safety
3 papers
LLM Safety Alignment is Divergence Estimation in Disguise
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons