2025
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
NeurIPS 2025spotlight
Reinforcement Learning algorithms commonly sample multiple ($n>1$) solution attempts for each problem and reward them independently. This optimizes for pass@1 performance and prioritizes individual sample performance over the diversity and collective utility of a set of samples. Such algorithms unde…