RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
authentic mathematical tasksautomated evaluationbenchmarkscompetition problemscontamination riskscontinually refreshable datasetdiverse research-level contentexperimental resultsformal proofshandling research mathematicsmathematical reasoningresearch environmentsvaluable assistantsverifiable statements
Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions