2026
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
CVPR 2026
Multimodal autoregressive (AR) models, based on next-token prediction and transformer architecture, have demonstrated remarkable capabilities in various multimodal tasks including text-to-image (T2I) generation. Despite their strong performance in general T2I tasks, our research reveals that these m