← Search

Yaowenqi Liu

1 accepted papers

2026

Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling

ICLR 2026poster

Process reward models (PRMs) are a cornerstone of test-time scaling (TTS), designed to verify and select the best responses from large language models (LLMs). However, this promise is challenged by recent benchmarks where simple majority voting, which ignores PRM signals, occasionally outperforms st…

Cited by 0SourceScholar