NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
harmful prompts
4 papers
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
Reasoning as an Adaptive Defense for Safety
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment