2026
Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
ICLR 2026poster
Process reward models (PRMs) are a cornerstone of test-time scaling (TTS), designed to verify and select the best responses from large language models (LLMs). However, this promise is challenged by recent benchmarks where simple majority voting, which ignores PRM signals, occasionally outperforms st…