2026
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
ICRA 2026poster
Vision-language models (VLMs) excel in visual understanding but often lack reliable grounding capabilities and actionable inference rates. Integrating them with open-vocabulary object detection (OVD), instance segmentation, and tracking leverages their strengths while mitigating these drawbacks. We …