NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
evaluation methodology
3 papers
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
More of the Same: Persistent Representational Harms Under Increased Representation