AAAI 2026technical0 citations

Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity

Guangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang, Tian Bai

Abstract

Vision-language object tracking overcomes the limitations of relying solely on visual features by leveraging language descriptions of objects to provide cross-modal semantic information, thereby enhancing model robustness in complex scenarios. However, most existing high-performance vision-language trackers are trained jointly on pure visual data and vision-language multimodal data. Due to the relative sparsity of language annotations in the data, the trackers tend to prioritize the localization role of visual features, diminishing the model

BibTeX
@inproceedings{aaai2026_awaredistillatio,
  title = {Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity},
  author = {Guangtong Zhang and Bineng Zhong and Shirui Yang and Yang Wang and Tian Bai},
  booktitle = {AAAI 2026},
  year = {2026}
}
Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity · AAAI 2026