ICML 2026poster0 citations

Value-as-Return: A Two-Stage Framework to Align on the Optimal Score Function

Shikun Sun, Shuo Huang, Yiding Chen, Wen Sun, Jia Jia

Abstract

Reinforcement learning with diffusion models has shown strong potential, but existing approaches such as variants of Direct Preference Optimization (DPO) often rely on an inaccurate simplification: they equate trajectory likelihoods with final-state probabilities. This mismatch leads to suboptimal alignment. We address this limitation with a principled framework that leverages the optimal value function as the return for short trajectory segments. Our approach follows a two-stage procedure: (i) learning a value-distribution function to estimate segment-level returns, and (ii) applying our VRPO to refine the score function. We prove that, under sufficient model capacity, the resulting model is equivalent to training a diffusion process on the tilted distribution proportional to $p(x)\exp(\eta r(x))$. Experiments on large-scale diffusion models validate our analysis and show stable and consistent improvements over prior methods.

DiffusionRLOptimizationRetrieval
BibTeX
@inproceedings{
sun2026valueasreturn,
title={Value-as-Return: A Two-Stage Framework to Align on the Optimal Score Function},
author={Shikun Sun and Shuo Huang and Yiding Chen and Wen Sun and Jia Jia},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=V3DCFHEC5C}
}