NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
ai agents
4 papers
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks