← Search

Jiahuan Zhou

50 accepted papers

2026

Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning

ICML 2026poster

Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concerns regarding privacy and storage. Thus, generative replay has emerged as a viable alternative, synthesizing old data usin…

Cited by 0SourceScholar
2026

CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification

AAAI 2026technical

Lifelong person Re-IDentification (LReID) aims to match the same person employing continuously collected individual data from different scenarios. To achieve continuous all-day person matching across day and night, Visible-Infrared Lifelong person Re-IDentification (VI-LReID) focuses on sequential t

Cited by 0SourcePDFScholar
2026

COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs

CVPR 2026

Despite Multimodal Large Language Models (MLLMs) having shown impressive capabilities, they may suffer from hallucinations. Empirically, we find that MLLMs attend disproportionately to task-irrelevant background regions compared with text-only LLMs, implying spurious background-answer correlations.

Cited by 0SourceScholar
2026

Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models

AAAI 2026technical

Test-time prompt tuning for vision-language models has demonstrated impressive generalization capabilities under zero-shot settings. However, tuning the learnable prompts solely based on unlabeled test data may induce prompt optimization bias, ultimately leading to suboptimal performance on downstre

Cited by 0SourcePDFScholar
2026

Hyper-LLaVA: Hyperbolic Uncertainty-aware Modality-Balanced Routing for Multimodal Continual Instruction Tuning

ICML 2026poster

Multimodal Continual Instruction Tuning (MCIT) aims to exploit the incrementally accumulated knowledge to process multimodal inputs of diverse tasks, where parameter routing is an important technology. Existing advanced methods typically rely on sample to task center similarity and cross-modal fusio…

Cited by 0SourceScholar
2026

Learnability-Driven Knowledge Assimilation for Class-Incremental Semantic Segmentation

ICML 2026poster

Class-incremental semantic segmentation learns new classes while retaining old ones without access to past data. Although existing methods alleviate catastrophic forgetting on old classes, new-class performance remains limited. We identify the key bottleneck arises from low-margin regions, where the…

Cited by 0SourceScholar
2026

MSCD-GS: Motion-Separated Cooperative Deblurring Dynamic Reconstruction via Gaussian Splatting

CVPR 2026

Although 4D reconstruction based on Gaussian Splatting has achieved many impressive results, reconstructing real-world images captured by a casual monocular camera remains a significant challenge. In dynamic scenes, as the camera and objects move during the exposure time, these input images inevitab

Cited by 0SourceScholar
2026

Naming to Learn: Class Incremental Learning for Vision-Language Model with Unlabeled Data

ICLR 2026poster

Class Incremental Learning (CIL) enables models to adapt to evolving data distributions by learning new classes over time without revisiting previous data. While recent methods utilizing pre-trained models have shown promising results, they often assume access to fully labeled data for each incremen…

Cited by 0SourceScholar
2026

On the Plasticity and Stability for Post-Training Large Language Models

ICML 2026poster

Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability retention. We identify a root cause as the geometric conflict between plasticity and stability gradients, which leads t…

Cited by 0SourceScholar
2026

RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation

CVPR 2026

Recently, state space models have demonstrated efficient video segmentation through linear-complexity state space compression. However, Video Semantic Segmentation (VSS) requires pixel-level spatiotemporal modeling capabilities to maintain temporal consistency in segmentation of semantic objects. Wh

Cited by 0SourcecodeScholar
2026

Test-Time Perturbation Tuning with Delayed Feedback for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object pose. We attribute this brittleness to trajectory overfitting, where VLAs over-attend to the spurious correlation betwe

Cited by 0SourcecodeScholar
2026

Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification

CVPR 2026

Lifelong person re-identification (LReID) aims to learn from varying domains to obtain a unified person retrieval model. Existing LReID approaches typically focus on learning from scratch or a visual classification-pretrained model, while the Vision-Language Model (VLM) has shown generalizable knowl

Cited by 0SourcecodeScholar
2025

C$^2$Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning

NeurIPS 2025poster

Federated continual learning (FCL) tackles scenarios of learning from continuously emerging task data across distributed clients, where the key challenge lies in addressing both temporal forgetting over time and spatial forgetting simultaneously. Recently, prompt-based FCL methods have shown advance…

Cited by 0SourceScholar
2025

CAPrompt: Cyclic Prompt Aggregation for Pre-Trained Model Based Class Incremental Learning

AAAI 2025technical

Recently, prompt tuning methods for pre-trained models have demonstrated promising performance in Class Incremental Learning (CIL). These methods typically involve learning task-specific prompts and predicting the task ID to select the appropriate prompts for inference. However, inaccurate task ID p…

2025

Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time Adaptation

NeurIPS 2025poster

Continual Test-Time Adaptation (CTTA) aims to quickly fine-tune the model during the test phase so that it can adapt to multiple unknown downstream domain distributions without pre-acquiring downstream domain data. To this end, existing advanced CTTA methods mainly reduce the catastrophic forgettin…

Cited by 0SourcecodeScholar
2025

Componential Prompt-Knowledge Alignment for Domain Incremental Learning

ICML 2025poster

Domain Incremental Learning (DIL) aims to learn from non-stationary data streams across domains while retaining and utilizing past knowledge. Although prompt-based methods effectively store multi-domain knowledge in prompt parameters and obtain advanced performance through cross-domain prompt fusion…

2025

DASK: Distribution Rehearsing via Adaptive Style Kernel Learning for Exemplar-Free Lifelong Person Re-Identification

AAAI 2025technical

Lifelong person re-identification (LReID) is an important but challenging task that suffers from catastrophic forgetting due to significant domain gaps between training steps. Existing LReID approaches typically rely on data replay and knowledge distillation to mitigate this issue. However, data rep…

2025

DKC: Differentiated Knowledge Consolidation for Cloth-Hybrid Lifelong Person Re-identification

CVPR 2025poster

Lifelong person re-identification (LReID) aims to match the same person using sequentially collected data. However, due to the long-term nature of lifelong learning, the inevitable changes in human clothes prevent the model from relying on unified discriminative information (e.g., clothing style) to…

2025

DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes

AAAI 2025technical

Vision-centric autonomous driving systems require diverse data for robust training and evaluation, which can be augmented by manipulating object positions and appearances within existing scene captures. While recent advancements in diffusion models have shown promise in video editing, their applicat…

2025

GAPrompt: Geometry-Aware Point Cloud Prompt for 3D Vision Model

ICML 2025poster

Pre-trained 3D vision models have gained significant attention for their promising performance on point cloud data. However, fully fine-tuning these models for downstream tasks is computationally expensive and storage-intensive. Existing parameter-efficient fine-tuning (PEFT) approaches, which focus…

2025

High-dimension Prototype is a Better Incremental Object Detection Learner

ICLR 2025poster

Incremental object detection (IOD), surpassing simple classification, requires the simultaneous overcoming of catastrophic forgetting in both recognition and localization tasks, primarily due to the significantly higher feature space complexity. Integrating Knowledge Distillation (KD) would mitigate…

Cited by 0SourcePDFScholar
2025

SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Prompting

CVPR 2025poster

Vision-language models (VLMs) encounter considerable challenges when adapting to domain shifts stemming from changes in data distribution. Test-time adaptation (TTA) has emerged as a promising approach to enhance VLM performance under such conditions. In practice, test data often arrives in batches,…

2025

STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding

CVPR 2025poster

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging due to limited labeled video data and high training costs. Rec…

2025

Selective Visual Prompting in Vision Mamba

AAAI 2025technical

Pre-trained Vision Mamba~(Vim) models have demonstrated exceptional performance across various computer vision tasks in a computationally efficient manner, attributed to their unique design of selective state space models. To further extend their applicability to diverse downstream vision tasks, Vim…

2025

Self-Reinforcing Prototype Evolution with Dual-Knowledge Cooperation for Semi-Supervised Lifelong Person Re-Identification

ICCV 2025poster

Current lifelong person re-identification (LReID) methods predominantly rely on fully labeled data streams. However, in real-world scenarios where annotation resources are limited, a vast amount of unlabeled data coexists with scarce labeled samples, leading to the Semi-Supervised LReID (Semi-LReID)…

2025

State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding

NeurIPS 2025poster

Recently, pre-trained state space models have shown great potential for video classification, which sequentially compresses visual tokens in videos with linear complexity, thereby improving the processing efficiency of video data while maintaining high performance. To apply powerful pre-trained mode…

Cited by 0SourceScholar
2025

Token Coordinated Prompt Attention is Needed for Visual Prompting

ICML 2025poster

Visual prompting techniques are widely used to efficiently fine-tune pretrained Vision Transformers (ViT) by learning a small set of shared prompts for all tokens. However, existing methods overlook the unique roles of different tokens in conveying discriminative information and interact with all to…

2025

UPP: Unified Point-Level Prompting for Robust Point Cloud Analysis

ICCV 2025poster

Pre-trained point cloud analysis models have shown promising advancements in various downstream tasks, yet their effectiveness is typically suffering from low-quality point cloud (i.e., noise and incompleteness), which is a common issue in real scenarios due to casual object occlusions and unsatisfa…

Cited by 0SourcePDFScholar
2024

Continual Vision-Language Retrieval via Dynamic Knowledge Rectification

AAAI 2024technical

The recent large-scale pre-trained models like CLIP have aroused great concern in vision-language tasks. However, when required to match image-text data collected in a streaming manner, namely Continual Vision-Language Retrieval (CVRL), their performances are still limited due to the catastrophic fo…

2024

DART: Dual-Modal Adaptive Online Prompting and Knowledge Retention for Test-Time Adaptation

AAAI 2024technical

As an up-and-coming area, CLIP-based pre-trained vision-language models can readily facilitate downstream tasks through the zero-shot or few-shot fine-tuning manners. However, they still face critical challenges in test-time generalization due to the shifts between the training and test data distrib…

Cited by 10SourcePDFScholar
2024

Distribution-aware Knowledge Prototyping for Non-exemplar Lifelong Person Re-identification

CVPR 2024poster

Lifelong person re-identification (LReID) suffers from the catastrophic forgetting problem when learning from non-stationary data. Existing exemplar-based and knowledge distillation-based LReID methods encounter data privacy and limited acquisition capacity respectively. In this paper we instead int…

2024

FCS: Feature Calibration and Separation for Non-Exemplar Class Incremental Learning

CVPR 2024poster

Non-Exemplar Class Incremental Learning (NECIL) involves learning a classification model on a sequence of data without access to exemplars from previously encountered old classes. Such a stringent constraint always leads to catastrophic forgetting of the learned knowledge. Currently existing methods…

2024

FashionERN: Enhance-and-Refine Network for Composed Fashion Image Retrieval

AAAI 2024technical

The goal of composed fashion image retrieval is to locate a target image based on a reference image and modified text. Recent methods utilize symmetric encoders (e.g., CLIP) pre-trained on large-scale non-fashion datasets. However, the input for this task exhibits an asymmetric nature, where the ref…

Cited by 5SourcePDFScholar
2024

FineFMPL: Fine-grained Feature Mining Prompt Learning for Few-Shot Class Incremental Learning

IJCAI 2024poster

Few-Shot Class Incremental Learning (FSCIL) aims to continually learn new classes with few training samples without forgetting already learned old classes. Existing FSCIL methods generally fix the backbone network in incremental sessions to achieve a balance between suppressing forgetting old classe…

2024

LSTKC: Long Short-Term Knowledge Consolidation for Lifelong Person Re-identification

AAAI 2024technical

Lifelong person re-identification (LReID) aims to train a unified model from diverse data sources step by step. The severe domain gaps between different training steps result in catastrophic forgetting in LReID, and existing methods mainly rely on data replay and knowledge distillation techniques to…

Cited by 13SourcePDFScholar
2024

Learning Continual Compatible Representation for Re-indexing Free Lifelong Person Re-identification

CVPR 2024poster

Lifelong Person Re-identification (L-ReID) aims to learn from sequentially collected data to match a person across different scenes. Once an L-ReID model is updated using new data all historical images in the gallery are required to be re-calculated to obtain new features for testing known as "re-in…

2024

Make Lossy Compression Meaningful for Low-Light Images

AAAI 2024technical

Low-light images frequently occur due to unavoidable environmental influences or technical limitations, such as insufficient lighting or limited exposure time. To achieve better visibility for visual perception, low-light image enhancement is usually adopted. Besides, lossy image compression is vita…

2024

SNIDA: Unlocking Few-Shot Object Detection with Non-linear Semantic Decoupling Augmentation

CVPR 2024poster

Once only a few-shot annotated samples are available the performance of learning-based object detection would be heavily dropped. Many few-shot object detection (FSOD) methods have been proposed to tackle this issue by adopting image-level augmentations in linear manners. Nevertheless those handcraf…

Cited by 9SourcePDFScholar
2023

Store and Fetch Immediately: Everything Is All You Need for Space-Time Video Super-resolution

AAAI 2023technical

Existing space-time video super-resolution (ST-VSR) methods fail to achieve high-quality reconstruction since they fail to fully explore the spatial-temporal correlations, long-range components in particular. Although the recurrent structure for ST-VSR adopts bidirectional propagation to aggregate i…

2022

Balancing between Forgetting and Acquisition in Incremental Subpopulation Learning

ECCV 2022poster

"The subpopulation shifting challenge, known as some subpopulations of a category that are not seen during training, severely limits the classification performance of the state-of-the-art convolutional neural networks. Thus, to mitigate this practical issue, we explore incremental subpopulation lear…

2020

Online Joint Multi-Metric Adaptation From Frequent Sharing-Subset Mining for Person Re-Identification

CVPR 2020poster

Person Re-IDentification (P-RID), as an instance-level recognition problem, still remains challenging in computer vision community. Many P-RID works aim to learn faithful and discriminative features/metrics from offline training data and directly use them for the unseen online testing data. However,…

Cited by 61PDFScholar
2020

Uncertainty-Aware Score Distribution Learning for Action Quality Assessment

CVPR 2020oral

Assessing action quality from videos has attracted growing attention in recent years. Most existing approaches usually tackle this problem based on regression algorithms, which ignore the intrinsic ambiguity in the score labels caused by multiple judges or their subjective appraisals. To address thi…

Cited by 171PDFcodeScholar
2019

Learning Robust Facial Landmark Detection via Hierarchical Structured Ensemble

ICCV 2019poster

Heatmap regression-based models have significantly advanced the progress of facial landmark detection. However, the lack of structural constraints always generates inaccurate heatmaps resulting in poor landmark detection performance. While hierarchical structure modeling methods have been proposed t…

Cited by 80PDFScholar
2018

Easy Identification From Better Constraints: Multi-Shot Person Re-Identification From Reference Constraints

CVPR 2018poster

Multi-shot person re-identification (MsP-RID) utilizes multiple images from the same person to facilitate identification. Considering the fact that motion information may not be discriminative nor reliable enough for MsP-RID, this paper is focused on handling the large variations in the visual appea…

Cited by 21SourcePDFScholar
2017

Efficient Online Local Metric Adaptation via Negative Samples for Person Re-Identification

ICCV 2017poster

Many existing person re-identification (PRID) methods typically attempt to train a faithful global metric offline to cover the enormous visual appearance variations, so as to directly use it online on various probes for identity matching. However, their need for a huge set of positive training pairs…

Cited by 98PDFScholar