2026
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
AAAI 2026technical
Visual grounding, the task of linking textual queries to specific regions within images, plays a pivotal role in vision-language integration. Existing methods typically rely on extensive task-specific annotations and fine-tuning, limiting their ability to generalize effectively to novel or out-of-di