AAAI 2026technical0 citations
Benchmarking LLMs’ Mathematical Reasoning with Unseen Random Variables Questions
Zijin Hong, Hao Wu, Su Dong, Junnan Dong, Yilin Xiao, Yujing Zhang, Zhu Wang, Feiran Huang
Abstract
Recent studies have raised significant concerns regarding the reliability of current mathematical benchmarks, highlighting key limitations such as simplistic design and potential data contamination that undermine evaluation accuracy. Consequently, developing a reliable benchmark that effectively evaluates large language models
BibTeX
@inproceedings{aaai2026_benchmarkingllms,
title = {Benchmarking LLMs’ Mathematical Reasoning with Unseen Random Variables Questions},
author = {Zijin Hong and Hao Wu and Su Dong and Junnan Dong and Yilin Xiao and Yujing Zhang and Zhu Wang and Feiran Huang and Linyi Li and Hongxia Yang and Xiao Huang},
booktitle = {AAAI 2026},
year = {2026}
}