← Search

Xiaotang Chen

10 accepted papers

2025

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking

ICCV 2025poster

Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, i…

2025

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features

ICML 2025poster

Effectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the mod…

2025

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

ICASSP 2025accepted

Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalit…

Cited by 0SourceScholar
2024

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts

NeurIPS 2024poster

Vision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed…

Cited by 5SourcePDFScholar
2022

Learning Disentangled Attribute Representations for Robust Pedestrian Attribute Recognition

AAAI 2022technical

Although various methods have been proposed for pedestrian attribute recognition, most studies follow the same feature learning mechanism, ie, learning a shared pedestrian image feature to classify multiple attributes. However, this mechanism leads to low-confidence predictions and non-robustness of…

Cited by 39SourcePDFScholar
2021

Spatial and Semantic Consistency Regularizations for Pedestrian Attribute Recognition

ICCV 2021poster

While recent studies on pedestrian attribute recognition have shown remarkable progress in leveraging complicated networks and attention mechanisms, most of them neglect the inter-image relations and an important prior: spatial consistency and semantic consistency of attributes under surveillance sc…

Cited by 79PDFScholar
2019

Towards Rich Feature Discovery With Class Activation Maps Augmentation for Person Re-Identification

CVPR 2019poster

The fundamental challenge of small inter-person variation requires Person Re-Identification (Re-ID) models to capture sufficient fine-grained information. This paper proposes to discover diverse discriminative visual cues without extra assistance, e.g., pose estimation, human parsing. Specifically,…

Cited by 311PDFScholar
2018

Adversarially Occluded Samples for Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is the task of retrieving particular persons across different cameras. Despite its great progress in recent years, it is still confronted with challenges like pose variation, occlusion, and similar appearance among different persons. The large gap between training and…

Cited by 305SourcePDFScholar
2017

Beyond Triplet Loss: A Deep Quadruplet Network for Person Re-Identification

CVPR 2017spotlight

Person re-identification (ReID) is an important task in wide area video surveillance which focuses on identifying people across different cameras. Recently, deep learning networks with a triplet loss become a common framework for person ReID. However, the triplet loss pays main attentions on obtaini…

Cited by 1545PDFScholar
2017

Learning Deep Context-Aware Features Over Body and Latent Parts for Person Re-Identification

CVPR 2017poster

Person Re-identification (ReID) is to identify the same person across different cameras. It is a challenging task due to the large variations in person pose, occlusion, background clutter, etc. How to extract powerful features is a fundamental problem in ReID and is still an open problem today. In t…

Cited by 829PDFScholar