2026
Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
ICLR 2026poster
Optimizing discrete diffusion model (DDM) with rewards remains a challenge—the non-autoregressive paradigm makes importance sampling intractable and rollout complex, puzzling reinforcement learning methods such as Group Relative Policy Optimization (GRPO). In this study, we introduce **MaskGRPO**, t…