ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, i…