NeurIPS 2025poster0 citations

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

Yifu Luo, Xinhao Hu, Keyu Fan, Haoyuan Sun, Zeyu Chen, Bo Xia, Tiantian Zhang, Yongzhe Chang

Abstract

Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO, the first method to incorporate Group Relative Policy Optimization (GRPO)-based RL into this overlooked paradigm. Our core insight is to redefine the transition probability, which is different from current approaches, and formulate the unmasking process as a multi-step decision-making problem. To further enhance our method, we explore several useful strategies, including removing the Kullback–Leibler constraint, applying the reduction strategy, and filtering out low-quality samples. Using Mask-GRPO, we improve a base model, Show-o, with substantial improvements on standard T2I benchmarks and preference alignment, outperforming existing state-of-the-art approaches.

reinforcement learningmasked autoregressive modelstext-to-image model
BibTeX
@inproceedings{
luo2025reinforcement,
title={Reinforcement Learning Meets Masked Generative Models: Mask-{GRPO} for Text-to-Image Generation},
author={Yifu Luo and Xinhao Hu and Keyu Fan and Haoyuan Sun and Zeyu Chen and Bo Xia and Tiantian Zhang and Yongzhe Chang and Xueqian Wang},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=C2QMbkp7iq}
}
Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation · NeurIPS 2025