← Search

Bineng Zhong

33 accepted papers

2026

An Efficient Token Compression Framework for Visual Object Tracking

CVPR 2026

Refining visual representations by eliminating their internal feature-level redundancy is crucial for simultaneously optimizing the performance and computational cost of models in visual tracking. To enhance their performance, many contemporary Transformer-based trackers leverage a larger number of

Cited by 0SourcecodeScholar
2026

Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity

AAAI 2026technical

Vision-language object tracking overcomes the limitations of relying solely on visual features by leveraging language descriptions of objects to provide cross-modal semantic information, thereby enhancing model robustness in complex scenarios. However, most existing high-performance vision-language

Cited by 0SourcePDFScholar
2026

Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning

CVPR 2026

Learning robust contextual knowledge from unlabeled videos is essential for advancing self-supervised tracking. However, conventional self-supervised trackers lack effective context modeling, while existing context association methods based on non-semantic queries struggle to adapt to unlabeled trac

Cited by 0SourceScholar
2026

Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking

CVPR 2026

Given the real-time demands of UAV tracking, many methods simplify the backbone to reduce computation, but this often weakens feature representation and degrades performance in complex scenarios. To alleviate this issue, we propose EATrack, an efficient and asymmetric UAV tracking framework centered

Cited by 0SourceScholar
2026

Learning to Track Instance from Single Nature Language Description

CVPR 2026

How to achieve vision-language (VL) tracking using natural language descriptions from a video sequence without relying on any bounding-box ground truth? In this work, we achieve this goal by tackling self-supervised VL tracking, which aims to evaluate tracking capabilities guided by natural language

Cited by 0SourceScholar
2026

MUTrack: A Memory-Aware Unified Representation Framework for Visual Tracking

AAAI 2026technical

Building a unified target representation that simultaneously achieves short-term adaptability and long-term stability is crucial for robust visual tracking. However, existing trackers typically face an inherent trade-off. Methods primarily relying on short-term appearance and motion cues achieve ra

Cited by 0SourcePDFScholar
2026

Motion-Aware Object Tracking via Motion and Geometry-Aware Cues

AAAI 2026technical

Understanding motion is essential for visual object tracking, especially in complex and dynamic scenarios. Yet, many existing methods rely on simplistic strategies such as template updates or temporal feature propagation, often overlooking the deeper modeling of motion information. To mitigate this

Cited by 0SourcePDFScholar
2026

Toward Low-Cost yet Effective Temporal Learning for UAV Tracking

CVPR 2026

The utilization of temporal information has always been an open topic in the tracking community. However, existing trackers tend to employ more and more inputs or parameters for temporal learning, hindering their deployment in resource-constrained unmanned aerial vehicles (UAVs). More importantly, t

Cited by 0SourcecodeScholar
2025

Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking

AAAI 2025technical

The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale and diversity of existing tracking datasets. In this work, we present a novel Self-Supervised Tracking framework, named S…

2025

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking

AAAI 2025technical

Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal re…

2025

Less Is More: Token Context-Aware Learning for Object Tracking

AAAI 2025technical

Recently, several studies have shown that utilizing contextual information to perceive target states is crucial for object tracking. They typically capture context by incorporating multiple video frames. However, these naive frame-context methods fail to consider the importance of each patch within…

2025

MambaLCT: Boosting Tracking via Long-term Context State Space Model

AAAI 2025technical

Effectively constructing context information with long-term dependencies from video sequences is crucial for object tracking. However, the context length constructed by existing work is limited, only considering object information from adjacent frames or video clips, leading to insufficient utilizat…

2025

Robust Tracking via Mamba-based Context-aware Token Learning

AAAI 2025technical

How to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming learning that combining temporal and appearance information by input more and more images (or features). Consequently, t…

2025

Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking

CVPR 2025poster

Vision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV) tracking which extremely emphasizes efficiency. In this study, we discover that many layers within lightweight ViT-bas…

2024

Autoregressive Queries for Adaptive Tracking with Spatio-Temporal Transformers

CVPR 2024poster

The rich spatio-temporal information is crucial to capture the complicated target appearance variations in visual tracking. However most top-performing tracking algorithms rely on many hand-crafted components for spatio-temporal information aggregation. Consequently the spatio-temporal information i…

2024

DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-based Human Video Generation

CVPR 2024poster

Existing diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In this paper we propose a novel framework DiffPerformer to synthesize high-fidelity an…

Cited by 1SourcePDFScholar
2024

Diffusion Mask-Driven Visual-language Tracking

IJCAI 2024poster

Most existing visual-language trackers greatly rely on the initial language descriptions on a target object to extract their multi-modal features. However, the initial language descriptions are often inaccurate in a highly time-varying video sequence and thus greatly deteriorate their tracking perfo…

Cited by 2SourcePDFScholar
2024

Explicit Visual Prompts for Visual Object Tracking

AAAI 2024technical

How to effectively exploit spatio-temporal information is crucial to capture target appearance changes in visual tracking. However, most deep learning-based trackers mainly focus on designing a complicated appearance model or template updating strategy, while lacking the exploitation of context betw…

2024

Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring

CVPR 2024poster

Existing joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models they typically rely on the assumed degra…

2024

ODTrack: Online Dense Temporal Token Learning for Visual Tracking

AAAI 2024technical

Online contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, t…

2023

Interactive Object Placement with Reinforcement Learning

ICML 2023poster

Object placement aims to insert a foreground object into a background image with a suitable location and size to create a natural composition. To predict a diverse distribution of placements, existing methods usually establish a one-to-one mapping from random vectors to the placements. However, thes…

Cited by 6SourcePDFScholar
2021

Discover Cross-Modality Nuances for Visible-Infrared Person Re-Identification

CVPR 2021poster

Visible-infrared person re-identification (Re-ID) aims to match the pedestrian images of the same identity from different modalities. Existing works mainly focus on alleviating the modality discrepancy by aligning the distributions of features from different modalities. However, nuanced but discrimi…

Cited by 286PDFcodeScholar
2021

Distractor-Aware Fast Tracking via Dynamic Convolutions and MOT Philosophy

CVPR 2021poster

A practical long-term tracker typically contains three key properties, i.e., an efficient model design, an effective global re-detection strategy and a robust distractor awareness mechanism. However, most state-of-the-art long-term trackers (e.g., Pseudo and re-detecting based ones) do not take all…

Cited by 53PDFcodeScholar
2021

EC-DARTS: Inducing Equalized and Consistent Optimization Into DARTS

ICCV 2021poster

Based on the relaxed search space, differential architecture search (DARTS) is efficient in searching for a high-performance architecture. However, the unbalanced competition among operations that have different trainable parameters causes the model collapse. Besides, the inconsistent structures in…

Cited by 9PDFcodeScholar
2021

Learning To Filter: Siamese Relation Network for Robust Tracking

CVPR 2021poster

Despite the great success of Siamese-based trackers, their performance under complicated scenarios is still not satisfying, especially when there are distractors. To this end, we propose a novel Siamese relation network, which introduces two efficient modules, i.e. Relation Detector (RD) and Refinem…

Cited by 146PDFcodeScholar
2020

Projection & Probability-Driven Black-Box Attack

CVPR 2020poster

Generating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer from the need for excessive queries, as it is non-trivial to find an appropriate direction to optimize in the high-dimens…

Cited by 58PDFcodeScholar
2020

Siamese Box Adaptive Network for Visual Tracking

CVPR 2020poster

Most of the existing trackers usually rely on either a multi-scale searching scheme or pre-defined anchor boxes to accurately estimate the scale and aspect ratio of a target. Unfortunately, they typically call for tedious and heuristic configurations. To address this issue, we propose a simple yet e…

Cited by 1031PDFcodeScholar
2020

What Can Be Transferred: Unsupervised Domain Adaptation for Endoscopic Lesions Segmentation

CVPR 2020poster

Unsupervised domain adaptation has attracted growing research attention on semantic segmentation. However, 1) most existing models cannot be directly applied into lesions transfer of medical images, due to the diverse appearances of same lesion among different datasets; 2) equal attention has been p…

Cited by 173PDFScholar
2019

Residual Non-local Attention Networks for Image Restoration

ICLR 2019poster

In this paper, we propose a residual non-local attention network for high-quality image restoration. Without considering the uneven distribution of information in the corrupted images, previous methods are restricted by local convolutional operation and equal treatment of spatial- and channel-wise f…

2018

Image Super-Resolution Using Very Deep Residual Channel Attention Networks

ECCV 2018poster

Convolutional neural network (CNN) depth is of crucial importance for image super-resolution (SR). However, we observe that deeper networks for image SR are more difficult to train. The low-resolution (LR) inputs and features contain abundant low-frequency information, which is treated equally acros…

2018

Residual Dense Network for Image Super-Resolution

CVPR 2018poster

In this paper, we propose dense feature fusion (DFF) for image super-resolution (SR). As the same content in different natural images often have various scales and angles of view, jointly leaning hierarchical features is essential for image SR. On the other hand, very deep convolutional neural netwo…

2015

Understanding Image Structure via Hierarchical Shape Parsing

CVPR 2015poster

Exploring image structure is a long-standing yet important research subject in the computer vision community. In this paper, we focus on understanding image structure inspired by the "simple-to-complex" biological evidence. A hierarchical shape parsing strategy is proposed to partition and organize…

Cited by 12SourcePDFScholar