NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
accuracy assessment
4 papers
How Benchmark Prediction from Fewer Data Misses the Mark
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents