2024
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
ECCV 2024poster
"Vision-language (VL) models often exhibit a limited understanding of complex expressions of visual objects (, attributes, shapes, and their relations), given complex and diverse language queries. Traditional approaches attempt to improve VL models using hard negative synthetic text, but their effec…