2025
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
ICCV 2025poster
Text-to-image (T2I) diffusion models excel at generating photorealistic images, but commonly struggle to render accurate spatial relationships described in text prompts. We identify two core issues underlying this common failure: 1) the ambiguous nature of spatial-related data in existing datasets,…