← Search

Yaozong Zheng

7 accepted papers

2026

Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning

CVPR 2026

Learning robust contextual knowledge from unlabeled videos is essential for advancing self-supervised tracking. However, conventional self-supervised trackers lack effective context modeling, while existing context association methods based on non-semantic queries struggle to adapt to unlabeled trac

Cited by 0SourceScholar
2026

Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking

CVPR 2026

Given the real-time demands of UAV tracking, many methods simplify the backbone to reduce computation, but this often weakens feature representation and degrades performance in complex scenarios. To alleviate this issue, we propose EATrack, an efficient and asymmetric UAV tracking framework centered

Cited by 0SourceScholar
2026

Learning to Track Instance from Single Nature Language Description

CVPR 2026

How to achieve vision-language (VL) tracking using natural language descriptions from a video sequence without relying on any bounding-box ground truth? In this work, we achieve this goal by tackling self-supervised VL tracking, which aims to evaluate tracking capabilities guided by natural language

Cited by 0SourceScholar
2025

Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking

AAAI 2025technical

The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale and diversity of existing tracking datasets. In this work, we present a novel Self-Supervised Tracking framework, named S…

2025

Less Is More: Token Context-Aware Learning for Object Tracking

AAAI 2025technical

Recently, several studies have shown that utilizing contextual information to perceive target states is crucial for object tracking. They typically capture context by incorporating multiple video frames. However, these naive frame-context methods fail to consider the importance of each patch within…

2025

Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking

CVPR 2025poster

Vision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV) tracking which extremely emphasizes efficiency. In this study, we discover that many layers within lightweight ViT-bas…

2024

ODTrack: Online Dense Temporal Token Learning for Visual Tracking

AAAI 2024technical

Online contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, t…