2026
ConsistCompose: Unified Multimodal Layout Control for Image Composition
CVPR 2026
Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding--aligning language with image regions--while their generative counterpart, linguistic-embedded layout-grounded generation(LELG) for layout-controll