← Search

Simon Matrenok

1 accepted papers

2025

Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions

NeurIPS 2025poster

Aligning large language models with pointwise absolute rewards has so far required online, on-policy algorithms such as PPO and GRPO. In contrast, simpler methods that can leverage offline or off-policy data, such as DPO and REBEL, are limited to learning from preference pairs or relative signals. T…

Cited by 0SourceScholar