IJCAI 20260 citations

MathCritique: Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Zhiheng Xi, Dingwen Yang, Jixuan Huang, Jiafu Tang, Xin Guo, Guanyu Li, Yiwen Ding, Wei He

Abstract

Training critique models to provide useful feedback for actor models is an effective approach in scalable oversight, especially for complex tasks like math reasoning. However, current research lacks suitable datasets for effectively training critique models and integrating them in a principled way at both test time and training time. To bridge this gap, we first propose AutoMathCritique, an automated and scalable framework for collecting critique data. Using it, we create MathCritique-76k, a dataset of $76,321$ instances paired with step-level feedback, which is then used to train critique models for generating natural-language feedback. We show that critique models consistently improve the actor’s performance—particularly on difficult queries at test time—with larger gains as inference compute scales. Building on the findings, we propose a critique-in-the-loop self-improvement method that incorporates critique-based supervision into the actor’s self-training process. Extensive experiments demonstrate that this method improves the actor’s exploration efficiency and solution diversity, especially on challenging queries, leading to a stronger actor model. Our code and datasets are at https://mathcritique.github.io/.

Natural Language Processing: Language generationNatural Language Processing: Language models
BibTeX
@inproceedings{ijcai2026_mathcritiqueenha,
  title = {MathCritique: Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision},
  author = {Zhiheng Xi and Dingwen Yang and Jixuan Huang and Jiafu Tang and Xin Guo and Guanyu Li and Yiwen Ding and Wei He and Boyang Hong and Shihan Dou and WenYu Zhan and Xiao Wang and Xiaowei Shi and Yitao Zhai and Rongxiang Weng and Jingang Wang and Rui Zheng and Tao Ji and Tao Gui and Zuxuan Wu and Qi Zhang and Xipeng Qiu and Xuanjing Huang and Yu-Gang Jiang},
  booktitle = {IJCAI 2026},
  year = {2026}
}
MathCritique: Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision · IJCAI 2026