NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
harmful requests
3 papers
Reasoning as an Adaptive Defense for Safety
The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
Understanding and Rectifying Safety Perception Distortion in VLMs