2024
Textual Tokens Classification for Multi-Modal Alignment in Vision-Language Tracking
ICASSP 2024accepted
Most vision-language (VL) trackers rely on coarse-grained information from sentences to achieve multi-modal alignment. However, this information is insufficient for accurately describing the target in each frame due to the inherent ambiguity, summarization, and invariance of sentences, thereby makin…