NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
safety performance
3 papers
Faithful Dynamic Imitation Learning from Human Intervention with Dynamic Regret Minimization
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
Training-Free Safe Denoisers for Safe Use of Diffusion Models