ACL 2025finding0 citations

The Lessons of Developing Process Reward Models in Mathematical Reasoning

Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jingren Zhou

Abstract

Process Reward Models (PRMs) aim to identify and mitigate intermediate errors in the reasoning processes in mathematical reasoning of Large Language Models (LLMs).However, the development of effective PRMs faces significant challenges, particularly in data annotation and evaluation methodologies.In this paper, through extensive experiments, we demonstrate that commonly used Monte Carlo (MC) estimation-based data synthesis for PRMs typically yields inferior performance and generalization compared to LLM-as-a-judge and human annotation methods.Furthermore, we identify potential biases in conventional Best-of-N (BoN) evaluation strategies for PRMs.To address these challenges, we develop a consensus filtering mechanism that effectively integrates MC estimation with LLM-as-a-judge and advocates a more comprehensive evaluation framework that combines response-level and step-level metrics. Based on the mechanisms, we significantly improve both model performance and data efficiency in the BoN evaluation and the step-wise error identification task.Finally, we release a new state-of-the-art PRM that outperforms existing open-source alternatives and provides practical guidelines for future research.

BibTeX
@inproceedings{zhang-etal-2025-lessons,
    title = "The Lessons of Developing Process Reward Models in Mathematical Reasoning",
    author = "Zhang, Zhenru  and
      Zheng, Chujie  and
      Wu, Yangzhen  and
      Zhang, Beichen  and
      Lin, Runji  and
      Yu, Bowen  and
      Liu, Dayiheng  and
      Zhou, Jingren  and
      Lin, Junyang",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.547/",
    doi = "10.18653/v1/2025.findings-acl.547",
    pages = "10495--10516",
    ISBN = "979-8-89176-256-5"
}
The Lessons of Developing Process Reward Models in Mathematical Reasoning · ACL 2025