2025
From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
ACL 2025long
Inference-time alignment methods have gained significant attention for their efficiency and effectiveness in aligning large language models (LLMs) with human preferences. However, existing dominant approaches using reward-guided search (RGS) primarily rely on outcome reward models (ORMs), which suff…