2026
Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
ICLR 2026poster
Process Reward Models (PRMs) have emerged as a promising approach to enhance the reasoning capabilities of large language models (LLMs) by guiding their step-by-step reasoning toward a final answer. However, existing PRMs either treat each reasoning step in isolation, failing to capture inter-step d…