← Search

Bojian Xiong

1 accepted papers

2025

HighMATH: Evaluating Math Reasoning of Large Language Models in Breadth and Depth

EMNLP 2025

With the rapid development of large language models (LLMs) in math reasoning, the accuracy of models on existing math benchmarks has gradually approached 90% or even higher. More challenging math benchmarks are hence urgently in need to satisfy the increasing evaluation demands. To bridge this gap,