2025
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
ICML 2025oral
Process Reward Models (PRMs) have proven effective at enhancing mathematical reasoning for Large Language Models (LLMs) by leveraging increased inference-time computation. However, they are predominantly trained on mathematical data and their generalizability to non-mathematical domains has not been…