From Curiosity to Caution: Mitigating Reward Hacking for Best-of-$N$ with Pessimism
Inference-time compute scaling has emerged as a powerful paradigm for improving language model performance on a wide range of tasks, but the question of how best to use the additional compute remains open. A popular approach is *Best-of-$N$* (BoN) sampling, where $N$ candidate responses are generat…