NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
safety evaluation
3 papers
AgentAuditor: Human-level Safety and Security Evaluation for LLM Agents
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks