2025
RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
NeurIPS 2025poster
Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions---failing to capture the nature of mathematics encountered in actual research environments. We introduce \textsc{Real…