← Search

Guangtong Zhang

3 accepted papers

2026

Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity

AAAI 2026technical

Vision-language object tracking overcomes the limitations of relying solely on visual features by leveraging language descriptions of objects to provide cross-modal semantic information, thereby enhancing model robustness in complex scenarios. However, most existing high-performance vision-language

Cited by 0SourcePDFScholar
2024

Diffusion Mask-Driven Visual-language Tracking

IJCAI 2024poster

Most existing visual-language trackers greatly rely on the initial language descriptions on a target object to extract their multi-modal features. However, the initial language descriptions are often inaccurate in a highly time-varying video sequence and thus greatly deteriorate their tracking perfo…

Cited by 2SourcePDFScholar