2026
iGRPO: Fast Online RL for Flow Matching Model with Dense Reward
ICML 2026poster
Conventional practice assumes that online reinforcement learning for flow-matching models requires sampling full denoising trajectories to compute rewards. This assumption underlies methods such as Group Relative Policy Optimization (GRPO), where the policy must traverse the entire reverse process b…