NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Thomas Foster
3 papers
University of OxfordFAIR @
LILO: Learning to Reason at the Frontier of Learnability
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements