2026
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
CVPR 2026
Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image alignment, often failing to faithfully depict objects with correct attributes such as color, shape, and spatial relation