2026
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
ICML 2026poster
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward,…