← Search

Wenhan Yang

69 accepted papers

2026

Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven Consistency

AAAI 2026technical

Super-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applicatio

Cited by 0SourcePDFScholar
2026

Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural Networks

ICLR 2026poster

Spiking neural networks (SNNs) compute with discrete spikes and exploit temporal structure, yet most adversarial attacks change intensities or event counts instead of timing. We study a timing-only adversary that retimes existing spikes while preserving spike counts and amplitudes in event-driven SN…

Cited by 0SourcecodeScholar
2026

Towards Generalized Representations for Low-Light Understanding: When Signal Constancy Meets Semantic Enrichment

CVPR 2026

Low-light degradation hampers machine understanding at night. Existing methods either overfit labeled data (paired supervision) or specific distributions (unpaired supervision), resulting in poor generalization under unseen degradations. In this paper, we propose UniPrior, a unified prior-based low-

Cited by 0SourceScholar
2026

Which Reasoning Traces Are Worth Generating Further? Data Curation for Training Reasoning Models

ICML 2026poster

Supervised fine-tuning (SFT) on a small high-quality set of long reasoning traces is an effective way to enable strong reasoning abilities for Large Language Models (LLMs). However, curating a high-quality SFT data requires generating a large pool of long Chain of Thoughts (CoTs), and filtering the …

Cited by 0SourceScholar
2025

Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time Shifts

ICCV 2025poster

Accurate monocular 3D object detection (M3OD) is pivotal for safety-critical applications like autonomous driving, yet its reliability deteriorates significantly under real-world domain shifts caused by environmental or sensor variations. To address these shifts, Test-Time Adaptation (TTA) methods h…

2025

Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable Trigger

AAAI 2025technical

No-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are suscep…

2025

End-to-End Low-Light Enhancement for Object Detection with Learned Metadata from RAWs

NeurIPS 2025poster

Although RAW images offer advantages over sRGB by avoiding ISP-induced distortion and preserving more information in low-light conditions, their widespread use is limited due to high storage costs, transmission burdens, and the need for significant architectural changes for downstream tasks. To addr…

Cited by 0SourceScholar
2025

Fast Omni-Directional Image Super-Resolution: Adapting the Implicit Image Function with Pixel and Semantic-Wise Spherical Geometric Priors

AAAI 2025technical

In the context of Omni-Directional Image (ODI) Super-Resolution (SR), the unique challenge arises from the non-uniform oversampling characteristics caused by EquiRectangular Projection (ERP). Considerable efforts in designing complex spherical convolutions or polyhedron reprojection offer significan…

2025

MTL-UE: Learning to Learn Nothing for Multi-Task Learning

ICML 2025poster

Most existing unlearnable strategies focus on preventing unauthorized users from training single-task learning (STL) models with personal data. Nevertheless, the paradigm has recently shifted towards multi-task data and multi-task learning (MTL), targeting generalist and foundation models that can h…

2025

Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures

ICLR 2025poster

Training with larger mini-batches improves the convergence rate and can yield superior performance. However, training with large mini-batches becomes prohibitive for Large Language Models (LLMs), due to the large GPU memory requirement. To address this problem, an effective approach is finding small…

Cited by 0SourcePDFScholar
2025

PRE-Mamba: A 4D State Space Model for Ultra-High-Frequent Event Camera Deraining

ICCV 2025poster

Event cameras excel in high temporal resolution and dynamic range but suffer from dense noise in rainy conditions. Existing event deraining methods face trade-offs between temporal precision, deraining effectiveness, and computational efficiency. In this paper, we propose PRE-Mamba, a novel point-ba…

2025

Privacy-Shielded Image Compression: Defending Against Exploitation from Vision-Language Pretrained Models

ICML 2025poster

The improved semantic understanding of vision-language pretrained (VLP) models has made it increasingly difficult to protect publicly posted images from being exploited by search engines and other similar tools. In this context, this paper seeks to protect users' privacy by implementing defenses at…

Cited by 0SourcePDFScholar
2025

Prompt-Guided Alignment with Information Bottleneck Makes Image Compression Also a Restorer

NeurIPS 2025poster

Learned Image Compression (LIC) models face critical challenges in real-world scenarios due to various environmental degradations, such as fog and rain. Due to the distribution mismatch between degraded inputs and clean training data, well-trained LIC models suffer from reduced compression efficienc…

Cited by 0SourceScholar
2025

Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems

CVPR 2025highlight

By locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve p…

2025

UP-Restorer: When Unrolling Meets Prompts for Unified Image Restoration

AAAI 2025technical

All-in-one restoration needs to implicitly distinguish between different degradation conditions and apply specific prior constraints accordingly. To fulfill this goal, our work makes the first effort to create an all-in-one restoration via unrolling from the typical maximum a-posterior optimization…

Cited by 0SourcePDFScholar
2025

Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision

AAAI 2025technical

The image compression model has long struggled with adaptability and generalization, as the decoded bitstream typically serves only human or machine needs and fails to preserve information for unseen visual tasks. Therefore, this paper innovatively introduces supervision obtained from multimodal pre…

2025

Which Tasks Should Be Compressed Together? A Causal Discovery Approach for Efficient Multi-Task Representation Compression

ICLR 2025poster

Conventional image compression methods are inadequate for intelligent analysis, as they overemphasize pixel-level precision while neglecting semantic significance and the interaction among multiple tasks. This paper introduces a Taskonomy-Aware Multi-Task Compression framework comprising (1) inter-…

Cited by 0SourcePDFScholar
2024

A Unified Image Compression Method for Human Perception and Multiple Vision Tasks

ECCV 2024poster

"Recent advancements in end-to-end image compression demonstrate the potential to surpass traditional codecs regarding rate-distortion performance. However, current methods either prioritize human perceptual quality or solely optimize for one or a few predetermined downstream tasks, neglecting a mor…

Cited by 0SourcePDFScholar
2024

Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks

ICML 2024poster

Contrastive Language-Image Pre-training (CLIP) on large image-caption datasets has achieved remarkable success in zero-shot classification and enabled transferability to new domains. However, CLIP is extremely more vulnerable to targeted data poisoning and backdoor attacks compared to supervised lea…

2024

ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance Field

AAAI 2024technical

Neural Radiance Fields (NeRF) have demonstrated impressive potential in synthesizing novel views from dense input, however, their effectiveness is challenged when dealing with sparse input. Existing approaches that incorporate additional depth or semantic supervision can alleviate this issue to an e…

2024

ContextGS : Compact 3D Gaussian Splatting with Anchor Level Context Model

NeurIPS 2024poster

Recently, 3D Gaussian Splatting (3DGS) has become a promising framework for novel view synthesis, offering fast rendering speeds and high fidelity. However, the large number of Gaussians and their associated attributes require effective compression techniques. Existing methods primarily compress ne…

2024

Correcting Diffusion-Based Perceptual Image Compression with Privileged End-to-End Decoder

ICML 2024poster

The images produced by diffusion models can attain excellent perceptual quality. However, it is challenging for diffusion models to guarantee distortion, hence the integration of diffusion models and image compression models still needs more comprehensive explorations. This paper presents a diffusio…

Cited by 3SourcePDFScholar
2024

DDR: Exploiting Deep Degradation Response as Flexible Image Descriptor

NeurIPS 2024poster

Image deep features extracted by pre-trained networks are known to contain rich and informative representations. In this paper, we present Deep Degradation Response (DDR), a method to quantify changes in image deep features under varying degradation conditions. Specifically, our approach facilitates…

2024

DeS3: Adaptive Attention-Driven Self and Soft Shadow Removal Using ViT Similarity

AAAI 2024technical

Removing soft and self shadows that lack clear boundaries from a single image is still challenging. Self shadows are shadows that are cast on the object itself. Most existing methods rely on binary shadow masks, without considering the ambiguous boundaries of soft and self shadows. In this paper, we…

2024

E3M: Zero-Shot Spatio-Temporal Video Grounding with Expectation-Maximization Multimodal Modulation

ECCV 2024oral

"Spatio-temporal video grounding aims to localize the spatio-temporal tube in a video according to the given language query. To eliminate the annotation costs, we make a first exploration to tackle spatio-temporal video grounding in a zero-shot manner. Our method dispenses with the need for any trai…

2024

Image Coding for Analytics via Adversarially Augmented Adaptation

ICASSP 2024accepted

Image Coding for Machine (ICM) aims to compress an image so that the reconstructed one can meet the requirements of both human vision and machine vision. Existing methods apply the constraint from the downstream models to improve machine analytics performance while compromising the visual quality. T…

Cited by 0SourceScholar
2024

Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video Grounding

AAAI 2024technical

This paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently com…

Cited by 12SourcePDFScholar
2024

Misalignment-Robust Frequency Distribution Loss for Image Transformation

CVPR 2024poster

This paper aims to address a common challenge in deep learning-based image transformation methods such as image enhancement and super-resolution which heavily rely on precisely aligned paired datasets with pixel-level alignments. However creating precisely aligned paired images presents significant…

2024

Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets Consistency

AAAI 2024technical

Natural language video localization plays a pivotal role in video understanding, and leveraging weakly-labeled data is considered a promising approach to circumvent the laborintensive process of manual annotations. However, this approach encounters two significant challenges: 1) limited input distri…

Cited by 9SourcePDFScholar
2024

Purify Unlearnable Examples via Rate-Constrained Variational Autoencoders

ICML 2024poster

Unlearnable examples (UEs) seek to maximize testing error by making subtle modifications to training examples that are correctly labeled. Defenses against these poisoning attacks can be categorized based on whether specific interventions are adopted during training. The first approach is training-ti…

2024

Region-Adaptive Transform with Segmentation Prior for Image Compression

ECCV 2024poster

"Learned Image Compression (LIC) has shown remarkable progress in recent years. Existing works commonly employ CNN-based or Transformer-based modules as transform methods for compression. However, there is no prior research on neural transform that focuses on specific regions. In response, we introd…

2024

Seeing Dark Videos via Self-Learned Bottleneck Neural Representation

AAAI 2024technical

Enhancing low-light videos in a supervised style presents a set of challenges, including limited data diversity, misalignment, and the domain gap introduced through the dataset construction pipeline. Our paper tackles these challenges by constructing a self-learned enhancement approach that gets rid…

2024

SinSR: Diffusion-Based Image Super-Resolution in a Single Step

CVPR 2024poster

While super-resolution (SR) methods based on diffusion models exhibit promising results their practical application is hindered by the substantial number of required inference steps. Recent methods utilize the degraded images in the initial state thereby shortening the Markov chain. Nevertheless the…

2024

Solving Diffusion ODEs with Optimal Boundary Conditions for Better Image Super-Resolution

ICLR 2024poster

Diffusion models, as a kind of powerful generative model, have given impressive results on image super-resolution (SR) tasks. However, due to the randomness introduced in the reverse process of diffusion models, the performances of diffusion-based SR models are fluctuating at every time of sampling,…

Cited by 9SourcePDFScholar
2024

Transferable Adversarial Attacks on SAM and Its Downstream Models

NeurIPS 2024poster

The utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical applications, their open accessibility also poses threats of adverse usage. This paper, for the first time, explores th…

2024

Unrolled Decomposed Unpaired Learning for Controllable Low-Light Video Enhancement

ECCV 2024poster

"Obtaining pairs of low/normal-light videos, with motions, is more challenging than still images, which raises technical issues and poses the technical route of unpaired learning as a critical role. This paper makes endeavors in the direction of learning for low-light video enhancement without using…

2023

Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger

CVPR 2023poster

Recent deep-learning-based compression methods have achieved superior performance compared with traditional approaches. However, deep learning models have proven to be vulnerable to backdoor attacks, where some specific trigger patterns added to the input can lead to malicious behavior of the models…

Cited by 57SourcePDFScholar
2023

Boundary-Aware Divide and Conquer: A Diffusion-Based Solution for Unsupervised Shadow Removal

ICCV 2023poster

Recent deep learning methods have achieved superior results in shadow removal. However, most of these supervised methods rely on training over a huge amount of shadow and shadow-free image pairs, which require laborious annotations and may end up with poor model generalization. Shadows, in fact, onl…

Cited by 19PDFScholar
2023

Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event Localization

AAAI 2023technical

This paper for the first time explores audio-visual event localization in an unsupervised manner. Previous methods tackle this problem in a supervised setting and require segment-level or video-level event category ground-truth to train the model. However, building large-scale multi-modality dataset…

Cited by 9SourcePDFScholar
2023

Dual Prompt Learning for Continual Rain Removal from Single Images

IJCAI 2023poster

Recent efforts have achieved remarkable progress on single image deraining on the stationary distributed data. However, catastrophic forgetting raises practical concerns when applying these methods to real applications, where the data distributions change constantly. In this paper, we investigate th…

Cited by 2SourcePDFScholar
2023

Estimating Reflectance Layer from a Single Image: Integrating Reflectance Guidance and Shadow/Specular Aware Learning

AAAI 2023technical

Estimating the reflectance layer from a single image is a challenging task. It becomes more challenging when the input image contains shadows or specular highlights, which often render an inaccurate estimate of the reflectance layer. Therefore, we propose a two-stage learning method, including refle…

2023

ExposureDiffusion: Learning to Expose for Low-light Image Enhancement

ICCV 2023poster

Previous raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical distribution information, leading to visually undesirable results. Thi…

Cited by 65PDFcodeScholar
2023

Raw Image Reconstruction With Learned Compact Metadata

CVPR 2023poster

While raw images exhibit advantages over sRGB images (e.g. linearity and fine-grained quantization level), they are not widely used by common users due to the large storage requirements. Very recent works propose to compress raw images by designing the sampling masks in the raw image pixel space, le…

2023

Robust Contrastive Language-Image Pretraining against Data Poisoning and Backdoor Attacks

NeurIPS 2023poster

Contrastive vision-language representation learning has achieved state-of-the-art performance for zero-shot classification, by learning from millions of image-caption pairs crawled from the internet. However, the massive data that powers large multimodal models such as CLIP, makes them extremely vul…

2023

ShadowDiffusion: When Degradation Prior Meets Diffusion Model for Shadow Removal

CVPR 2023poster

Recent deep learning methods have achieved promising results in image shadow removal. However, their restored images still suffer from unsatisfactory boundary artifacts, due to the lack of degradation prior and the deficiency in modeling capacity. Our work addresses these issues by proposing a unifi…

2022

Low-Light Image Enhancement with Normalizing Flow

AAAI 2022technical

To enhance low-light images to normally-exposed ones is highly ill-posed, namely that the mapping relationship between them is one-to-many. Previous works based on the pixel-wise reconstruction losses and deterministic processes fail to capture the complex conditional distribution of normally expose…

2022

MSDN: Mutually Semantic Distillation Network for Zero-Shot Learning

CVPR 2022poster

The key challenge of zero-shot learning (ZSL) is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus achieving a desirable knowledge transfer to unseen classes. Prior works either simply align the global features of an image with its associated…

Cited by 177PDFcodeScholar
2022

Self-Learned Video Super-Resolution with Augmented Spatial and Temporal Context

ICASSP 2022accepted

Video super-resolution methods typically rely on paired training data, in which the low-resolution frames are usually synthetically generated under predetermined degradation conditions (e.g., Bicubic downsampling). However, in real applications, it is labor-consuming and expensive to obtain this kin…

Cited by 0SourceScholar
2022

Semantic Compression Embedding for Generative Zero-Shot Learning

IJCAI 2022poster

Generative methods have been successfully applied in zero-shot learning (ZSL) by learning an implicit mapping to alleviate the visual-semantic domain gaps and synthesizing unseen samples to handle the data imbalance between seen and unseen classes. However, existing generative methods simply use vis…

2022

Semantically Contrastive Learning for Low-Light Image Enhancement

AAAI 2022technical

Low-light image enhancement (LLE) remains challenging due to the unfavorable prevailing low-contrast and weak-visibility problems of single RGB images. In this paper, we respond to the intriguing learning-related question -- if leveraging both accessible unpaired over/underexposed images and high-le…

2022

Towards Robust Rain Removal Against Adversarial Attacks: A Comprehensive Benchmark Analysis and Beyond

CVPR 2022poster

Rain removal aims to remove rain streaks from images/videos and reduce the disruptive effects caused by rain. It not only enhances image/video visibility but also allows many computer vision algorithms to function properly. This paper makes the first attempt to conduct a comprehensive study on the r…

Cited by 66PDFcodeScholar
2022

URetinex-Net: Retinex-Based Deep Unfolding Network for Low-Light Image Enhancement

CVPR 2022poster

Retinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement. However, the commonly used hand-crafted priors and optimization-driven solutions lead to the absence of adaptivity and efficiency. To address these issues, in…

Cited by 612PDFcodeScholar
2022

Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects Suppression

ECCV 2022poster

"Night images suffer not only from low light, but also from uneven distributions of light. Most existing night visibility enhancement methods focus mainly on enhancing low-light regions. This inevitably leads to over enhancement and saturation in bright regions, such as those regions affected by lig…

2021

Self-Aligned Video Deraining With Transmission-Depth Consistency

CVPR 2021poster

In this paper, we address the problems of rain streaks and rain accumulation removal in video, by developing a self-aligned network with transmission-depth consistency. Existing video based deraining method focus only on rain streak removal, and commonly use optical flow to align the rain video fram…

Cited by 40PDFcodeScholar
2021

Teacher-Student Learning With Multi-Granularity Constraint Towards Compact Facial Feature Representation

ICASSP 2021accepted

In this paper, we propose a novel end-to-end feature compression scheme by leveraging the representation and learning capability of deep neural networks, towards intelligent front-end equipped analysis with promising accuracy and efficiency. In particular, the extracted features are compactly coded…

Cited by 0SourceScholar
2020

From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement

CVPR 2020poster

Under-exposure introduces a series of visual degradation, i.e. decreased visibility, intensive noise, and biased color, etc. To address these problems, we propose a novel semi-supervised learning approach for low-light image enhancement. A deep recursive band network (DRBN) is proposed to recover a…

Cited by 657PDFScholar
2020

Self-Learning Video Rain Streak Removal: When Cyclic Consistency Meets Temporal Correspondence

CVPR 2020poster

In this paper, we address the problem of rain streaks removal in video by developing a self-learned rain streak removal method, which does not require any clean groundtruth images in the training process. The method is inspired by fact that the adjacent frames are highly correlated and can be regard…

Cited by 82PDFcodeScholar
2018

Attentive Generative Adversarial Network for Raindrop Removal From a Single Image

CVPR 2018poster

Raindrops adhered to a glass window or camera lens can severely hamper the visibility of a background scene and degrade an image considerably. In this paper, we address the problem by visually removing raindrops, and thus transforming a raindrop degraded image into a clean one. The problem is intrac…

Cited by 851SourcePDFScholar
2018

Erase or Fill? Deep Joint Recurrent Rain Removal and Reconstruction in Videos

CVPR 2018poster

In this paper, we address the problem of video rain removal by constructing deep recurrent convolutional networks. We visit the rain removal case by considering rain occlusion regions, i.e. light transmittance of rain streaks is low. Different from additive rain streaks, in such rain occlusion regio…

Cited by 224SourcePDFScholar
2017

Deep Joint Rain Detection and Removal From a Single Image

CVPR 2017poster

In this paper, we address a rain removal problem from a single image, even in the presence of heavy rain and rain streak accumulation. Our core ideas lie in our new rain image model and new deep learning architecture. We add a binary map that provides rain streak locations to an existing model, whic…

Cited by 1366PDFScholar
2017

General scale interpolation via context-aware autoregressive model and multiplanar constraint

ICASSP 2017accepted

In this paper, we propose a novel image interpolation algorithm suitable for general scale enlargement. Different from previous AR-based interpolation algorithms which employ predetermined reference configuration to predict pixel values, we consider the context information when building AR models. O…

Cited by 0SourceScholar
2016

Human activity recognition based on weighted limb features

IROS 2016poster

Human activity recognition plays an important role in personal assistive robot, being able to recognize human activity and perform corresponding assistive action is a great challenges for personal assistive robot. Human body is an articulated system of rigid segments that can be divided into five pa…

Cited by 3SourceScholar
2015

Neighborhood regression for edge-preserving image super-resolution

ICASSP 2015accepted

There have been many proposed works on image super-resolution via employing different priors or external databases to enhance HR results. However, most of them do not work well on the reconstruction of high-frequency details of images, which are more sensitive for human vision system. Rather than re…

Cited by 0SourceScholar
2015

Novel autoregressive model based on adaptive window-extension and patch-geodesic distance for image interpolation

ICASSP 2015accepted

In this paper, we propose a novel autoregressive (AR) model based on the adaptive window and the patch-geodesic distance for the image interpolation. The model combines the information of inner/inter-patch correlation. To model the inner-patch correlation, we introduce a patch-geodesic distance simi…

Cited by 0SourceScholar