AAAI 2026technical0 citations

Benchmarking LLMs’ Mathematical Reasoning with Unseen Random Variables Questions

Zijin Hong, Hao Wu, Su Dong, Junnan Dong, Yilin Xiao, Yujing Zhang, Zhu Wang, Feiran Huang

Abstract

Recent studies have raised significant concerns regarding the reliability of current mathematical benchmarks, highlighting key limitations such as simplistic design and potential data contamination that undermine evaluation accuracy. Consequently, developing a reliable benchmark that effectively evaluates large language models

BibTeX
@inproceedings{aaai2026_benchmarkingllms,
  title = {Benchmarking LLMs’ Mathematical Reasoning with Unseen Random Variables Questions},
  author = {Zijin Hong and Hao Wu and Su Dong and Junnan Dong and Yilin Xiao and Yujing Zhang and Zhu Wang and Feiran Huang and Linyi Li and Hongxia Yang and Xiao Huang},
  booktitle = {AAAI 2026},
  year = {2026}
}