2024
ST-LDM: A Universal Framework for Text-Grounded Object Generation in Real Images
ECCV 2024poster
"We present a novel image editing scenario termed Text-grounded Object Generation (TOG), defined as generating a new object in the real image spatially conditioned by textual descriptions. Existing diffusion models exhibit limitations of spatial perception in complex real-world scenes, relying on ad…