2025
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
ACL 2025finding
While large language models (LLMs) have significantly advanced mathematical reasoning, Process Reward Models (PRMs) have been developed to evaluate the logical validity of reasoning steps. However, PRMs still struggle with out-of-distribution (OOD) challenges. This paper identifies the OOD issues in…