NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Ruibin Yuan
4 papers
Carnegie Mellon University
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
OmniBench: Towards The Future of Universal Omni-Language Models
SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines