← Search

Aihua Zheng

16 accepted papers

2026

Chain-of-Thought Guided Multi-Modal Object Re-Identification

CVPR 2026

With the rise of visual-language models, multi-modal ReID retrieves specific targets by integrating different spectra and textual descriptions. Existing methods merely adopt descriptive representation learning for image-text, ignoring the relationships among the intrinsic logical hierarchies of sema

Cited by 0SourcecodeScholar
2026

Dual-Teacher Interactive Knowledge Distillation Network for Text-to-Visible & Infrared Person Retrieval

AAAI 2026technical

Text-to-visible & infrared person retrieval aims to retrieve the corresponding visible (RGB) and thermal infrared (TIR) images given the text descriptions. Existing methods perform semantic decoupling by aligning RGB and TIR features separately to different attributes, thereby facilitating the align

Cited by 0SourcePDFScholar
2026

Progressive Multi-modal Knowledge Distillation for Multi-spectral Object Re-identification

AAAI 2026technical

In the field of multi-spectral object re-identification (ReID), multi-modal knowledge and modal-specific knowledge exhibit complementary advantages when handling hard samples, but existing methods rarely integrate this collaborative information. Knowledge distillation is a direct approach for trans

Cited by 0SourcePDFScholar
2026

ProxyTTT: Proxy-driven Test-Time Training for Multi-modal Re-identification

AAAI 2026technical

Multi-modal object re-identification (ReID) aims to retrieve specific targets by leveraging complementary cues from different sensing modalities. Despite recent progress, two key challenges remain: (1) the limited ability to jointly address both modality and viewpoint discrepancies, and (2) the diff

Cited by 0SourcePDFScholar
2026

Semantic-Driven Visual Progressive Refinement for Aerial-Ground Person ReID: A Challenging Large-Scale Benchmark

AAAI 2026technical

Aerial-Ground Person Re-IDentification (AGPReID) aims to extract identity-discriminative representations from heterogeneous perspectives across different platforms in complex real-world environments. However, existing methods primarily focus on visual appearance modeling and make insufficient use of

Cited by 0SourcePDFScholar
2025

DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification

AAAI 2025technical

Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by combining complementary information from multiple modalities. Existing multi-modal object ReID methods primarily focus on the fusion of heterogeneous features. However, they often overlook the dynamic quality changes in…

2025

Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery Detection

ICASSP 2025accepted

With the advancement of generative AI, distinguishing real and AI-generated faces in videos has become increasingly challenging. However, traditional methods struggle to capture local details and temporal dynamics simultaneously, making it difficult to achieve high detection accuracy while maintaini…

Cited by 0SourceScholar
2025

MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic Prompt

AAAI 2025technical

Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by utilizing complementary image information from different modalities. Recently, large-scale pre-trained models like CLIP have demonstrated impressive performance in traditional single-modal ReID tasks. However, they rema…

2025

UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification

NeurIPS 2025poster

Multi-modal object Re-IDentification (ReID) has gained considerable attention with the goal of retrieving specific targets across cameras using heterogeneous visual data sources. At present, multi-modal object ReID faces two core challenges: (1) learning robust features under fine-grained local nois…

Cited by 0SourcecodeScholar
2024

Day-Night Cross-domain Vehicle Re-identification

CVPR 2024poster

Previous advances in vehicle re-identification (ReID) are mostly reported under favorable lighting conditions while cross-day-and-night performance is neglected which greatly hinders the development of related traffic intelligence applications. This work instead develops a novel Day-Night Dual-domai…

2024

Heterogeneous Test-Time Training for Multi-Modal Person Re-identification

AAAI 2024technical

Multi-modal person re-identification (ReID) seeks to mitigate challenging lighting conditions by incorporating diverse modalities. Most existing multi-modal ReID methods concentrate on leveraging complementary multi-modal information via fusion or interaction. However, the relationships among hetero…

2024

Parallel Augmentation and Dual Enhancement for Occluded Person Re-Identification

ICASSP 2024accepted

Occluded person re-identification (Re-ID), the task of searching for the same person’s images in occluded environments, has attracted lots of attention in the past decades. Recent approaches concentrate on improving performance on occluded data by data/feature augmentation or using extra models to p…

Cited by 0SourceScholar
2022

Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identification

AAAI 2022technical

Multi-modal person Re-ID introduces more complementary information to assist the traditional Re-ID task. Existing multi-modal methods ignore the importance of modality-specific information in the feature fusion stage. To this end, we propose a novel method to boost modality-specific representations…

Cited by 51SourcePDFScholar
2020

Arbitrary Talking Face Generation via Attentional Audio-Visual Coherence Learning

IJCAI 2020poster

Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on either disentangling the information in a single image or l…

Cited by 0SourcePDFScholar