DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with Attributes
Recent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capa…