RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics

Jie Zhang (SPY Lab -- ETH Zurich) · Cezara Petrui (ETHZ - ETH Zurich) · Kristina Nikolić (ETH Zurich) · Florian Tramer (ETH Zurich)
authentic mathematical tasksautomated evaluationbenchmarkscompetition problemscontamination riskscontinually refreshable datasetdiverse research-level contentexperimental resultsformal proofshandling research mathematicsmathematical reasoningresearch environmentsvaluable assistantsverifiable statements

Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions