AAAI 2026technical0 citations
Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity
Guangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang, Tian Bai
Abstract
Vision-language object tracking overcomes the limitations of relying solely on visual features by leveraging language descriptions of objects to provide cross-modal semantic information, thereby enhancing model robustness in complex scenarios. However, most existing high-performance vision-language trackers are trained jointly on pure visual data and vision-language multimodal data. Due to the relative sparsity of language annotations in the data, the trackers tend to prioritize the localization role of visual features, diminishing the model
BibTeX
@inproceedings{aaai2026_awaredistillatio,
title = {Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity},
author = {Guangtong Zhang and Bineng Zhong and Shirui Yang and Yang Wang and Tian Bai},
booktitle = {AAAI 2026},
year = {2026}
}