NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
harmful queries
3 papers
MetaDefense: Defending Fine-tuning based Jailbreak Attack Before and During Generation
VMDT: Decoding the Trustworthiness of Video Foundation Models
Why Do Some Language Models Fake Alignment While Others Don't?