Hybrid Layout Control for Diffusion Transformer: Fewer Annotations, Superior Aesthetics
Text-to-image generation models often struggle to interpret spatially aware text prompts effectively. To overcome this, existing approaches typically require millions of high-quality semantic layout annotations consisting of bounding boxes and regional prompts. This paper shows that the large amount…