MathNet — Explore 30,000+ Olympiad Math Problems
math-benchmarkolympiad-mathmathematical-reasoningllm-evaluationretrieval
Abstraction: Large-scale multilingual Olympiad math benchmark for LLM reasoning and retrieval
Key points:
- MathNet: 30,676 expert-authored Olympiad problems with solutions spanning 47 countries, 17 languages, and two decades of competitions
- Three tasks: Problem Solving (solve outright), Math-Aware Retrieval (find equivalent problems), and RAG (retrieval-augmented solving)
- State-of-the-art solving scores: Gemini-3.1-Pro 78.4%, GPT-5 69.3% — strong but not saturated
- Retrieval is the larger gap: Recall@1 below 5% for every embedding model tested on finding mathematically equivalent problems
- RAG gains depend heavily on retrieval quality; DeepSeek-V3.2-Speciale achieves up to +12% with good retrieval, reaching highest benchmark scores
- First benchmark for evaluating mathematical problem retrieval; data and benchmark publicly released
Connections: Mit · Gemini · GPT-5 · Mathematical Reasoning · Retrieval Augmented Generation · LLM Benchmarks
Source: https://mathnet.mit.edu/