← Search

Jin Tang

33 accepted papers

2026

Chain-of-Thought Guided Multi-Modal Object Re-Identification

CVPR 2026

With the rise of visual-language models, multi-modal ReID retrieves specific targets by integrating different spectra and textual descriptions. Existing methods merely adopt descriptive representation learning for image-text, ignoring the relationships among the intrinsic logical hierarchies of sema

Cited by 0SourcecodeScholar
2026

Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark

CVPR 2026

Text-aerial person retrieval aims to identify targets in UAV-captured images from eyewitness descriptions, supporting intelligent transportation and public security applications. Compared to ground-view text-image person retrieval, UAV-captured images often suffer from degraded visual information du

Cited by 0SourcecodeScholar
2026

NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training

CVPR 2026

Neural operators have emerged as an efficient paradigm for solving PDEs, overcoming the limitations of traditional numerical methods and significantly improving computational efficiency. However, due to the diversity and complexity of PDE systems, existing neural operators typically rely on a single

Cited by 0SourcecodeScholar
2026

Progressive Multi-cue Alignment for Unaligned RGBT Tracking

CVPR 2026

Unaligned RGBT tracking aims to achieve robust target localization across spatially misaligned RGB and thermal infrared (TIR) videos, a crucial challenge for deploying RGBT tracking in real-world scenarios. Existing methods often calculate all cross-modal alignment parameters (i.e., spatial shift an

Cited by 0SourcecodeScholar
2026

Progressive Multi-modal Knowledge Distillation for Multi-spectral Object Re-identification

AAAI 2026technical

In the field of multi-spectral object re-identification (ReID), multi-modal knowledge and modal-specific knowledge exhibit complementary advantages when handling hard samples, but existing methods rarely integrate this collaborative information. Knowledge distillation is a direct approach for trans

Cited by 0SourcePDFScholar
2026

ProxyTTT: Proxy-driven Test-Time Training for Multi-modal Re-identification

AAAI 2026technical

Multi-modal object re-identification (ReID) aims to retrieve specific targets by leveraging complementary cues from different sensing modalities. Despite recent progress, two key challenges remain: (1) the limited ability to jointly address both modality and viewpoint discrepancies, and (2) the diff

Cited by 0SourcePDFScholar
2026

Semantic-Driven Visual Progressive Refinement for Aerial-Ground Person ReID: A Challenging Large-Scale Benchmark

AAAI 2026technical

Aerial-Ground Person Re-IDentification (AGPReID) aims to extract identity-discriminative representations from heterogeneous perspectives across different platforms in complex real-world environments. However, existing methods primarily focus on visual appearance modeling and make insufficient use of

Cited by 0SourcePDFScholar
2026

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

CVPR 2026

Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be

Cited by 0SourceScholar
2025

AlignCAPE: Support and Query Feature Aligning for Category-Agnostic Pose Estimation

IROS 2025

Recent advancements in category-agnostic pose estimation have focused on developing a unified model capable of localizing keypoint coordinates across arbitrary categories, which enables robots to accurately interact with diverse objects by understanding their poses. While existing methods predominan

Cited by 0SourceScholar
2025

CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset

CVPR 2025poster

X-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence that can significantly reduce diagnostic burdens and patient wait times. Despite significant progress, we believe that the task has reached a bottleneck due to the limited benchmark datasets and the existi…

2025

MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection

ICASSP 2025accepted

Sound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from heterogeneous dataset. In this paper, we introduce a novel dual-branch a…

Cited by 0SourceScholar
2025

RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba

AAAI 2025technical

Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal representation, due to large computational burden. To address this issue…

Cited by 2SourcePDFScholar
2025

Template-based Uncertainty Multimodal Fusion Network for RGBT Tracking

IJCAI 2025

RGBT tracking is to localize the predefined targets in video sequences by effectively leveraging the information from both visible light (RGB) and thermal infrared (TIR) modalities. However, the quality of different modalities changes dynamically in complex scenes, and effectively perceiving modal q

2025

UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification

NeurIPS 2025poster

Multi-modal object Re-IDentification (ReID) has gained considerable attention with the goal of retrieving specific targets across cameras using heterogeneous visual data sources. At present, multi-modal object ReID faces two core challenges: (1) learning robust features under fine-grained local nois…

Cited by 0SourcecodeScholar
2024

Cross-Covariate Gait Recognition: A Benchmark

AAAI 2024technical

Gait datasets are essential for gait research. However, this paper observes that present benchmarks, whether conventional constrained or emerging real-world datasets, fall short regarding covariate diversity. To bridge this gap, we undertake an arduous 20-month effort to collect a cross-covariate ga…

2024

Event Stream-based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel Baseline

CVPR 2024poster

Tracking with bio-inspired event cameras has garnered increasing interest in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The former incurs higher inference costs while the latter may be susceptible to the impa…

Cited by 43SourcePDFScholar
2024

Structural Information Guided Multimodal Pre-training for Vehicle-Centric Perception

AAAI 2024technical

Understanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-train models on large-scale classification datasets and then fine-tune them for specific downstream tasks. However, they neg…

2022

Attribute-Based Progressive Fusion Network for RGBT Tracking

AAAI 2022technical

RGBT tracking usually suffers from various challenge factors, such as fast motion, scale variation, illumination variation, thermal crossover and occlusion, to name a few. Existing works often study fusion models to solve all challenges simultaneously, and it requires fusion models complex enough an…

Cited by 171SourcePDFScholar
2022

Few-Shot Class-Incremental Learning via Entropy-Regularized Data-Free Replay

ECCV 2022poster

"Few-shot class-incremental learning (FSCIL) has been proposed aiming to enable a deep learning system to incrementally learn new classes with limited data. Recently, a pioneer claims that the commonly used replay-based method in class-incremental learning (CIL) is ineffective and thus not preferred…

2022

Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identification

AAAI 2022technical

Multi-modal person Re-ID introduces more complementary information to assist the traditional Re-ID task. Existing multi-modal methods ignore the importance of modality-specific information in the feature fusion stage. To this end, we propose a novel method to boost modality-specific representations…

Cited by 51SourcePDFScholar
2022

Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-Experts

NeurIPS 2022accept

In this paper, we tackle the problem of domain shift. Most existing methods perform training on multiple source domains using a single model, and the same trained model is used on all unseen target domains. Such solutions are sub-optimal as each target domain exhibits its own specialty, which is not…

2022

MetaFSCIL: A Meta-Learning Approach for Few-Shot Class Incremental Learning

CVPR 2022poster

In this paper, we tackle the problem of few-shot class incremental learning (FSCIL). FSCIL aims to incrementally learn new classes with only a few samples in each class. Most existing methods only consider the incremental steps at test time. The learning objective of these methods is often hand-engi…

Cited by 186PDFScholar
2021

Test-Time Fast Adaptation for Dynamic Scene Deblurring via Meta-Auxiliary Learning

CVPR 2021poster

In this paper, we tackle the problem of dynamic scene deblurring. Most existing deep end-to-end learning approaches adopt the same generic model for all unseen test images. These solutions are sub-optimal, as they fail to utilize the internal information within a specific image. On the other hand, a…

Cited by 106PDFScholar
2020

All at Once: Temporally Adaptive Multi-Frame Interpolation with Advanced Motion Modeling

ECCV 2020poster

Recent advances in high refresh rate displays as well as the increased interest in high rate of slow motion and frame up-conversion fuel the demand for efficient and cost-effective multi-frame video interpolation solutions. To that regard, inserting multiple frames between consecutive video frames a…

Cited by 79SourcePDFScholar
2018

Cross-Modal Ranking with Soft Consistency and Noisy Labels for Robust RGB-T Tracking

ECCV 2018poster

Due to the complementary benefits of visible (RGB) and thermal infrared (T) data, RGB-T object tracking attracts more and more attention recently for boosting the performance under adverse illumination conditions. Existing RGB-T tracking methods usually localize a target object with a bounding box,…

Cited by 163SourcePDFScholar
2018

SINT++: Robust Visual Tracking via Adversarial Positive Instance Generation

CVPR 2018poster

Existing visual trackers are easily disturbed by occlusion,blurandlargedeformation. Inthechallengesofocclusion, motion blur and large object deformation, the performance of existing visual trackers may be limited due to the followingissues: i)Adoptingthedensesamplingstrategyto generate positive exam…

Cited by 152SourcePDFScholar
2015

SOLD: Sub-Optimal Low-rank Decomposition for Efficient Video Segmentation

CVPR 2015poster

This paper investigates how to perform robust and efficient unsupervised video segmentation while suppressing the effects of data noises and/or corruptions. We propose a general algorithm, called Sub-Optimal Low-rank Decomposition (SOLD), which pursues the low-rank representation for video segmentat…

Cited by 54SourcePDFScholar