← Search

Chenggang Yan

41 accepted papers

2026

2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information Propagation

AAAI 2026technical

Despite recent progress in adapting State Space Models such as Mamba to vision tasks, their intrinsic 1D scanning mechanism imposes limitations when applied to inherently 2D-structured data like images. Existing adaptations, including VMamba and 2DMamba, either suffer from inconsistency between scan

Cited by 0SourcePDFScholar
2026

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models

CVPR 2026

Prompt learning has emerged as an efficient alternative to fine-tuning pre-trained vision-language models (VLMs).Despite its promise, current methods still struggle to maintain tail-class discriminability when adapting to class-imbalanced datasets. In this work, we propose cluster-aware neural colla

Cited by 0SourceScholar
2026

EMR-Diff: Edge-aware Multimodal Residual Diffusion Model for Hyperspectral Image Super-resolution

CVPR 2026

Hardware constraints make it challenging to simultaneously acquire hyperspectral images (HSIs) with both high spatial and high spectral resolutions. A promising solution is to fuse low-resolution HSI (LR-HSI) with high-resolution multispectral images (HR-MSI) to generate high-resolution HSI (HR-HSI)

Cited by 0SourcecodeScholar
2026

Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection

CVPR 2026

Domain adaptive object detection (DAOD) aims to generalize an object detector trained on a source domain to a target domain, where the domain gap degrades the adaptability. Recently, large-scale vision foundation models (VFMs), pretrained on web-scale datasets, exhibit such powerful generalization c

Cited by 0SourceScholar
2026

Forgetting Knowledge Localization and Isolation for Continual Forgetting of Pre-trained Vision Models

AAAI 2026technical

Continual forgetting task aims to continuously remove multiple target knowledge subsets from pre-trained models while maintaining the integrity of remaining knowledge. Existing methods suffer from both incomplete forgetting of target knowledge and unintended forgetting of indistinguishable remaining

Cited by 0SourcePDFScholar
2026

Temporal Calibrating and Distilling for Scene-Text Aware Text-Video Retrieval

AAAI 2026technical

Existing text-video retrieval methods mainly focus on singlemodal video content (i.e., visual entities), often overlooking heterogeneous scene text that is ubiquitous in human environments. Although scene text in videos provides finegrained semantics for cross-modal retrieval, effectively utilizing

Cited by 0SourcePDFScholar
2025

Frequency Dynamic Convolution for Dense Image Prediction

CVPR 2025poster

While Dynamic Convolution (DY-Conv) has shown promising performance by enabling adaptive weight selection through multiple parallel weights combined with an attention mechanism, the frequency response of these weights tends to exhibit high similarity, resulting in high parameter costs but limited ad…

2025

K-Buffers: A Plug-in Method for Enhancing Neural Fields with Multiple Buffers

IJCAI 2025

Neural fields are now the central focus of research in 3D vision and computer graphics. Existing methods mainly focus on various scene representations, such as neural points and 3D Gaussians. However, few works have studied the rendering process to enhance the neural fields. In this work, we propose

2025

Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain Adaptation

CVPR 2025poster

This paper explores the Class-Incremental Source-Free Unsupervised Domain Adaptation (CI-SFUDA) problem, where the unlabeled target data come incrementally without access to labeled source instances. This problem poses two challenges, the interference of similar source-class knowledge in target-clas…

Cited by 1SourcePDFScholar
2025

Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning

ICCV 2025poster

Audio-visual multi-task incremental learning aims to continuously learn from multiple audio-visual tasks without the need for joint training on all tasks. The challenge of the problem is how to preserve the old task knowledge while facilitating the learning of new task with previous experiences. To…

2025

Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing

CVPR 2025poster

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip. This task demands the model bridge character performances and complicated proso…

2025

Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change Captioning

AAAI 2025technical

Change captioning aims to describe the differences between two similar images using natural language, significantly aiding in understanding and monitoring changes. This challenging task requires a fine-grained understanding of subtle changes while resisting disturbances like viewpoint shifts and ill…

Cited by 0SourcePDFScholar
2025

SMTPD: A New Benchmark for Temporal Prediction of Social Media Popularity

CVPR 2025poster

Social media popularity prediction task aims to predict the popularity of posts on social media platforms, which has a positive driving effect on application scenarios such as content optimization, digital marketing and online advertising. Though many studies have made significant progress, few of t…

2025

SdalsNet: Self-Distilled Attention Localization and Shift Network for Unsupervised Camouflaged Object Detection

AAAI 2025technical

Unsupervised camouflaged object detection (UCOD) poses significant challenges, primarily attributed to the absence of human labels. Existing UCOD methodologies, leveraging attention mechanisms, often struggle to achieve precise localization of camouflaged objects. To overcome this limitation, we int…

2025

ThermalGaussian: Thermal 3D Gaussian Splatting

ICLR 2025poster

Thermography is especially valuable for the military and other users of surveillance cameras. Some recent methods based on Neural Radiance Fields (NeRF) are proposed to reconstruct the thermal scenes in 3D from a set of thermal and RGB images. However, unlike NeRF, 3D Gaussian splatting (3DGS) preva…

2024

CURE4Rec: A Benchmark for Recommendation Unlearning with Deeper Influence

NeurIPS 2024poster

With increasing privacy concerns in artificial intelligence, regulations have mandated the right to be forgotten, granting individuals the right to withdraw their data from models. Machine unlearning has emerged as a potential solution to enable selective forgetting in models, particularly in recomm…

2024

Context-aware Difference Distilling for Multi-change Captioning

ACL 2024long

Multi-change captioning aims to describe complex and coupled changes within an image pair in natural language. Compared with single-change captioning, this task requires the model to have higher-level cognition ability to reason an arbitrary number of changes. In this paper, we propose a novel conte…

2024

Coupled Confusion Correction: Learning from Crowds with Sparse Annotations

AAAI 2024technical

As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sourcing has been widely adopted to alleviate the cost of collecting labels, which also inevitably introduces label noise a…

2024

Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning

ECCV 2024poster

"Change captioning aims to succinctly describe the semantic change between a pair of similar images, while being immune to distractors (illumination and viewpoint changes). Under these distractors, unchanged objects often appear pseudo changes about location and scale, and certain objects might over…

2024

Improving Radiology Report Generation with D2-Net: When Diffusion Meets Discriminator

ICASSP 2024accepted

Radiology report generation (RRG) aims to automatically provide observations and insight into a patient’s condition based on radiology images, which is able to greatly reduce the workload of physicians on the premise of ensuring the quality of medical treatment. Existing works leverage the Transform…

Cited by 0SourceScholar
2024

Quad Bayer Joint Demosaicing and Denoising Based on Dual Encoder Network with Joint Residual Learning

AAAI 2024technical

The recent imaging technology Quad Bayer CFA brings better imaging PSNR and higher visual quality compared to traditional Bayer CFA, but also serious challenges for demosaicing and denoising during the ISP pipeline. In this paper, we propose a novel dual encoder network, namely DRNet, to achieve joi…

Cited by 12SourcePDFScholar
2024

Rethinking Boundary Discontinuity Problem for Oriented Object Detection

CVPR 2024poster

Oriented object detection has been developed rapidly in the past few years where rotation equivariance is crucial for detectors to predict rotated boxes. It is expected that the prediction can maintain the corresponding rotation when objects rotate but severe mutation in angular prediction is someti…

2024

StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

ACL 2024findings

Given a script, the challenge in Movie Dubbing (Visual Voice Cloning, V2C) is to generate speech that aligns well with the video in both time and emotion, based on the tone of a reference audio track. Existing state-of-the-art V2C models break the phonemes in the script according to the divisions be…

2023

Gaussian Label Distribution Learning for Spherical Image Object Detection

CVPR 2023poster

Spherical image object detection emerges in many applications from virtual reality to robotics and automatic driving, while many existing detectors use ln-norms loss for regression of spherical bounding boxes. There are two intrinsic flaws for ln-norms loss, i.e., independent optimization of paramet…

Cited by 9SourcePDFScholar
2023

Improving Dynamic HDR Imaging with Fusion Transformer

AAAI 2023technical

Reconstructing a High Dynamic Range (HDR) image from several Low Dynamic Range (LDR) images with different exposures is a challenging task, especially in the presence of camera and object motion. Though existing models using convolutional neural networks (CNNs) have made great progress, challenges s…

2023

Self-supervised Cross-view Representation Reconstruction for Change Captioning

ICCV 2023poster

Change captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruc…

Cited by 36PDFcodeScholar
2023

Sph2Pob: Boosting Object Detection on Spherical Images with Planar Oriented Boxes Methods

IJCAI 2023poster

Object detection on panoramic/spherical images has been developed rapidly in the past few years, where IoU-calculator is a fundamental part of various detector components, i.e. Label Assignment, Loss and NMS. Due to the low efficiency and non-differentiability of spherical Unbiased IoU, spherical ap…

2022

Gait Recognition in the Wild With Dense 3D Representations and a Benchmark

CVPR 2022poster

Existing studies for gait recognition are dominated by 2D representations like the silhouette or skeleton of the human body in constrained scenes. However, humans live and walk in the unconstrained 3D space, so projecting the 3D human body onto the 2D plane will discard a lot of crucial information…

Cited by 192PDFcodeScholar
2022

PANDORA: A Panoramic Detection Dataset for Object with Orientation

ECCV 2022poster

"Panoramic images have become increasingly popular as omnidirectional panoramic technology has advanced. Many datasets and works resort to object detection to better understand the content of the panoramic image. These datasets and detectors use a Bounding Field of View (BFoV) as a bounding box in p…

2022

Unbiased IoU for Spherical Image Object Detection

AAAI 2022technical

As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while…

Cited by 12SourcePDFScholar
2021

Automated Model Design and Benchmarking of Deep Learning Models for COVID-19 Detection with Chest CT Scans

AAAI 2021technical

The COVID-19 pandemic has spread globally for several months. Because its transmissibility and high pathogenicity seriously threaten people's lives, it is crucial to accurately and quickly detect COVID-19 infection. Many recent studies have shown that deep learning (DL) based solutions can help dete…

2021

Rˆ3Net:Relation-embedded Representation Reconstruction Network for Change Captioning

EMNLP 2021main

Change captioning is to use a natural language sentence to describe the fine-grained disagreement between two similar images. Viewpoint change is the most typical distractor in this task, because it changes the scale and location of the objects and overwhelms the representation of real change. In th…

2020

Real-World Automatic Makeup via Identity Preservation Makeup Net

IJCAI 2020poster

This paper focuses on the real-world automatic makeup problem. Given one non-makeup target image and one reference image, the automatic makeup is to generate one face image, which maintains the original identity with the makeup style in the reference image. In the real-world scenario, face makeup ta…

Cited by 0SourcePDFScholar
2020

Unsupervised Person Re-Identification via Softened Similarity Learning

CVPR 2020poster

Person re-identification (re-ID) is an important topic in computer vision. This paper studies the unsupervised setting of re-ID, which does not require any labeled information and thus is freely deployed to new scenarios. There are very few studies under this setting, and one of the best approach ti…

Cited by 346PDFScholar
2019

Approximated Oracle Filter Pruning for Destructive CNN Width Optimization

ICML 2019oral

It is not easy to design and run Convolutional Neural Networks (CNNs) due to: 1) finding the optimal number of filters (i.e., the width) at each layer is tricky, given an architecture; and 2) the computational intensity of CNNs impedes the deployment on computationally limited devices. Oracle Prunin…

2019

Social Relation Recognition From Videos via Multi-Scale Spatial-Temporal Reasoning

CVPR 2019poster

Discovering social relations, e.g., kinship, friendship, etc., from visual contents can make machines better interpret the behaviors and emotions of human beings. Existing studies mainly focus on recognizing social relations from still images while neglecting another important media--video. On one h…

Cited by 94PDFScholar