NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Weiran Xu
3 papers
Stanford University
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents