2026
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
ICLR 2026poster
Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time scaling (TTS). However, while most PRMs exhibit substantial gains in mathematical domains, the scarcity of domain-specific…