NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
quality control
3 papers
CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
Position: Benchmarking is Broken - Don't Let AI be Its Own Judge
Scaling Physical Reasoning with the PHYSICS Dataset