2024
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
ECCV 2024poster
"Recently, diffusion models have increasingly demonstrated their capabilities in vision understanding. By leveraging prompt-based learning to construct sentences, these models have shown proficiency in classification and visual grounding tasks. However, existing approaches primarily showcase their a…