← Search

Shizhou Zhang

19 accepted papers

2026

Attention Retention for Continual Learning with Vision Transformers

AAAI 2026technical

Continual learning (CL) empowers AI systems to progressively acquire knowledge from non-stationary data streams. However, catastrophic forgetting remains a critical challenge. In this work, we identify attention drift in Vision Transformers as a primary source of catastrophic forgetting, where the a

Cited by 0SourcePDFScholar
2026

Better Matching, Less Forgetting: A Quality-Guided Matcher for Transformer-based Incremental Object Detection

AAAI 2026technical

Incremental Object Detection (IOD) aims to continuously learn new object classes without forgetting previously learned ones. A persistent challenge is catastrophic forgetting, primarily attributed to background shift in conventional detectors. While pseudo-labeling mitigates this in dense detectors,

Cited by 0SourcePDFScholar
2026

Do Large Language Models Reason About Uncertainty Like Humans? A Benchmark on Hurricane Forecast Visualization Comprehension

AAAI 2026technical

Uncertainty visualizations, such as hurricane cones and ensemble tracks, are essential for risk communication but are often misinterpreted, leading to harmful decisions. As AI assistants like large language models (LLMs) increasingly support understanding of graphics and decision-making, they offer

Cited by 0SourcePDFScholar
2026

DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance

AAAI 2026technical

Infrared imaging plays a critical role in low-light and adverse weather conditions. However, due to the distinct characteristics of infrared images, existing foundation models such as Masked Autoencoder (MAE) trained on visible data perform suboptimal in infrared image interpretation tasks. To bridg

Cited by 0SourcePDFScholar
2026

Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation

CVPR 2026

Incremental Object Detection (IOD) enables AI systems to continuously acquire new object classes while preserving knowledge of previously learned ones, an ability essential for deployment in dynamic, real-world environments. Existing IOD methods typically rely on knowledge distillation to mitigate c

Cited by 0SourceScholar
2026

Knowing the Unknown: Interpretable Open-World Object Detection via Concept Decomposition Model

ICML 2026poster

Open-world object detection (OWOD) requires incrementally detecting known categories while reliably identifying unknown objects. Existing methods primarily focus on improving unknown recall, yet overlook interpretability, often leading to known–unknown confusion and reduced prediction reliability. T…

Cited by 0SourceScholar
2026

OrienPose: Orientation-Guided Novel View Synthesis for Single-Image Unseen Object Pose Estimation

CVPR 2026

Estimating the 3D pose of unseen objects from a single image remains a fundamental yet challenging problem in computer vision, especially under a CAD model-free setting.Pioneering attempts address this issue by matching templates generated through Novel View Synthesis (NVS), which essentially aims t

Cited by 0SourcecodeScholar
2026

Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection

ICML 2026poster

Accurately localizing and segmenting small targets in low signal-to-noise ratio (SNR) infrared sequences remains a challenging task. Since targets are often indistinguishable from the background in individual frames, existing methods, even when equipped with advanced foundation model and powerful in…

Cited by 0SourceScholar
2026

YOLO-IOD: Towards Real Time Incremental Object Detection

AAAI 2026technical

Current methodologies for incremental object detection (IOD) primarily rely on Faster R-CNN or DETR series detectors; however, these approaches do not accommodate the real-time YOLO detection frameworks. In this paper, we first identify three primary types of knowledge conflicts that contribute to c

Cited by 0SourcePDFScholar
2025

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing

AAAI 2025technical

Large multimodal language models (MLLMs) have revolutionized natural language processing and visual understanding, but often contain outdated or inaccurate information. Current multimodal knowledge editing evaluations are limited in scope and potentially biased, focusing on narrow tasks and failing…

2025

Demystifying Catastrophic Forgetting in Two-Stage Incremental Object Detector

ICML 2025poster

Catastrophic forgetting is a critical chanllenge for incremental object detection (IOD). Most existing methods treat the detector monolithically, relying on instance replay or knowledge distillation without analyzing component-specific forgetting. Through dissection of Faster R-CNN, we reveal a key…

Cited by 0SourcePDFScholar
2025

Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpening

CVPR 2025poster

Pansharpening aims at integrating complementary information from panchromatic and multispectral images. Available deep-learning based pansharpening methods typically perform exceptionally with particular satellite datasets. At the same time, it has been observed that these models also exhibit scene…

2025

Gradient Decomposition and Alignment for Incremental Object Detection

ICCV 2025poster

Incremental object detection (IOD) is crucial for enabling AI systems to continuously learn new object classes over time while retaining knowledge of previously learned categories, allowing model to adapt to dynamic environments without forgetting prior information.Existing IOD methods primarily emp…

2025

Revisiting Generative Replay for Class Incremental Object Detection

CVPR 2025poster

Generative replay has gained significant attention in class-incremental learning; however, its application to Class Incremental Object Detection (CIOD) remains limited due to the challenges in generating complex images with precise spatial arrangements. In this study, motivated by the observation th…

2025

Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual Learning

AAAI 2025technical

Visual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generato…

Cited by 0SourcePDFScholar
2024

Cross-Platform Video Person ReID: A New Benchmark Dataset and Adaptation Approach

ECCV 2024poster

"In this paper, we construct a large-scale benchmark dataset for Ground-to-Aerial Video-based person Re-Identification, named G2A-VReID, which comprises 185,907 images and 5,576 tracklets, featuring 2,788 distinct identities. To our knowledge, this is the first dataset for video ReID under Ground-to…

2024

Visual Prompt Tuning in Null Space for Continual Learning

NeurIPS 2024poster

Existing prompt-tuning methods have demonstrated impressive performances in continual learning (CL), by selecting and updating relevant prompts in the vision-transformer models. On the contrary, this paper aims to learn each task by tuning the prompts in the direction orthogonal to the subspace span…

2022

Dynamically Transformed Instance Normalization Network for Generalizable Person Re-identification

ECCV 2022poster

"Existing person re-identification methods often suffer significant performance degradation on unseen domains, which fuels interest in domain generalizable person re-identification (DG-PReID). As an effective technology to alleviate domain variance, the Instance Normalization (IN) has been widely em…

Cited by 51SourcePDFScholar
2019

Vehicle Re-Identification in Aerial Imagery: Dataset and Approach

ICCV 2019poster

In this work, we construct a large-scale dataset for vehicle re-identification (ReID), which contains 137k images of 13k vehicle instances captured by UAV-mounted cameras. To our knowledge, it is the largest UAV-based vehicle ReID dataset. To increase intra-class variation, each vehicle is captured…

Cited by 78PDFScholar