← Search

WenQi Ren

60 accepted papers

2026

Decision-Driven Orthogonal Learning with Complementary Feature Mining for Robust Synthetic Image Detection

AAAI 2026technical

The widespread and inconsistent compression applied by Online Social Networks severely degrades the performance of synthetic image detectors. We attribute this degradation to two main issues: 1) the model confuses forgery artifacts with compression artifacts, and 2) compression erodes crucial discri

Cited by 0SourcePDFScholar
2026

Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening

CVPR 2026

In this work, we propose a Multigrain-aware Semantic Prototype Scanning paradigm for pan-sharpening, built upon a high-order RWKV architecture and a tri-token prompting mechanism derived from semantic clustering. Specifically, our method contains three key components: 1) Multigrain-aware Semantic Pr

Cited by 0SourceScholar
2026

Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learning

CVPR 2026

Few-Shot Class-Incremental Learning (FSCIL) poses a critical challenge in machine learning, requiring models to continuously integrate novel classes with limited samples while preserving knowledge of previously seen classes. While existing FSCIL approaches have demonstrated promising results, they s

Cited by 0SourceScholar
2026

VIRUS: Injecting Persistent Cognitive Pathogens into Stateful Zero-Shot Object Navigation Agents

ICML 2026poster

Zero-Shot Object Navigation (ZSON) agents rely on continuously updated internal states to support long-horizon planning and decision-making. However, existing methods heavily depend on the observational outputs of vision-language models (VLMs) during state updates and lack explicit validation of per…

Cited by 0SourceScholar
2026

WEVSR: Adapting Video Diffusion Generators to Real-World Video Super‑Resolution with Wavelet-Enhanced VAE Encoder

ICML 2026poster

Recent advances in video diffusion models have demonstrated remarkable generative capability, yet adapting these large pretrained text-to-video (T2V) models to video super‑resolution (VSR) typically encounters challenges, such as artifacts introduced by complex degradations in real-world scenarios a…

Cited by 0SourceScholar
2025

Critical Forgetting-Based Multi-Scale Disentanglement for Deepfake Detection

AAAI 2025technical

Recent face forgery detection methods based on disentangled representation learning utilize paired images for cross-reconstruction, aiming to extract forgery-relevant attributes and forgery-irrelevant content. However, there still exist the following issues that may comprise the detector performance…

Cited by 0SourcePDFScholar
2025

Dual Prompting Image Restoration with Diffusion Transformers

CVPR 2025poster

Recent state-of-the-art image restoration methods mostly adopt latent diffusion models with U-Net backbones, yet still facing challenges in achieving high-quality restoration due to their limited capabilities. Diffusion transformers (DiTs), like SD3, are emerging as a promising alternative because o…

Cited by 1SourcePDFScholar
2025

ECSNN: Spiking Neural Networks for Efficient Exposure Correction in Endoscopy Imaging

ICASSP 2025accepted

The quality of endoscopic images is critical to the success of polyp segmentation, highlighting the need for accurate exposure correction in endoscopy. While traditional deep learning methods are effective, they demand substantial computational resources during inference. To address this, we propose…

Cited by 0SourceScholar
2025

Gaze Label Alignment: Alleviating Domain Shift for Gaze Estimation

AAAI 2025technical

Gaze estimation methods encounter significant performance deterioration when being evaluated across different domains, because of the domain gap between the testing and training data. Existing methods try to solve this issue by reducing the deviation of data distribution, however, they ignore the ex…

Cited by 1SourcePDFScholar
2025

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models

ICCV 2025poster

With the rapid advancement of multimodal large language models (MLLMs), concerns regarding their security have increasingly captured the attention of both academia and industry. Although MLLMs are vulnerable to jailbreak attacks, designing effective jailbreak attacks poses unique challenges, especia…

2025

LVPTrack: High Performance Domain Adaptive UAV Tracking with Label Aligned Visual Prompt Tuning

AAAI 2025technical

Visual object tracking is essentially crucial for unmanned aerial vehicles (UAVs). Despite the substantial progress, most of the existing UAV trackers are designed for well-conditioned daytime data, while for the scenarios in challenging weather condition, e.g. foggy or nighttime environment, the tr…

Cited by 0SourcePDFScholar
2025

RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution

AAAI 2025technical

Benefiting from their powerful generative capabilities, pretrained diffusion models have garnered significant attention for real-world image super-resolution (Real-SR). Existing diffusion-based SR approaches typically utilize semantic information from degraded images and restoration prompts to activ…

2025

Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining

CVPR 2025poster

Significant progress has been made in video restoration under rainy conditions over the past decade, largely propelled by advancements in deep learning. Nevertheless, existing methods that depend on paired data struggle to generalize effectively to real-world scenarios, primarily due to the disparit…

Cited by 0SourcePDFScholar
2025

UMDATrack: Unified Multi-Domain Adaptive Tracking Under Adverse Weather Conditions

ICCV 2025poster

Visual object tracking has gained promising progress in past decades. Most of the existing approaches focus on learning target representation in well-conditioned daytime data, while for the unconstrained real-world scenarios with adverse weather conditions, e.g. nighttime or foggy environment, the t…

2025

Ultra-High-Definition Dynamic Multi-Exposure Image Fusion via Infinite Pixel Learning

AAAI 2025technical

With the continuous improvement of device imaging resolution, the popularity of Ultra-High-Definition (UHD) images is increasing. Unfortunately, existing methods for fusing multi-exposure images in dynamic scenes are designed for low-resolution images, which makes them inefficient for generating hig…

Cited by 0SourcePDFScholar
2025

Unsupervised Diffusion-Based Degradation Modeling for Real-World Super-Resolution

AAAI 2025technical

Single image super-solution (SR) aims to restore a high-resolution (HR) image from a degraded low-resolution (LR) image. However, existing SR models still face a significant domain gap between synthetic and real-world datasets due to the mismatched degradation distributions, hindering SR models from…

2024

CountFormer: Multi-View Crowd Counting Transformer

ECCV 2024poster

"Multi-view counting (MVC) methods have shown their superiority over single-view counterparts, particularly in situations characterized by heavy occlusion and severe perspective distortions. However, hand-crafted heuristic features and identical camera layout requirements in conventional MVC methods…

2024

EnsIR: An Ensemble Algorithm for Image Restoration via Gaussian Mixture Models

NeurIPS 2024poster

Image restoration has experienced significant advancements due to the development of deep learning. Nevertheless, it encounters challenges related to ill-posed problems, resulting in deviations between single model predictions and ground-truths. Ensemble learning, as a powerful machine learning tech…

2024

Frequency Aware and Graph Fusion Network for Polyp Segmentation

ICASSP 2024accepted

Polyp segmentation plays a crucial role in the prevention of colon cancer. However, the diverse shapes of polyps and their similarity to normal areas in terms of color and texture make polyp segmentation a challenging task. Currently, most polyp segmentation methods solely focus on spatial domain fe…

Cited by 0SourceScholar
2024

Hybrid Frequency Modulation Network for Image Restoration

IJCAI 2024poster

Image restoration involves recovering a high-quality image from its corrupted counterpart. This paper presents an effective and efficient framework for image restoration, termed CSNet, based on ``channel + spatial" hybrid frequency modulation. Different feature channels include different degradation…

2024

Insert or Attach: Taxonomy Completion via Box Embedding

ACL 2024long

Taxonomy completion, enriching existing taxonomies by inserting new concepts as parents or attaching them as children, has gained significant interest. Previous approaches embed concepts as vectors in Euclidean space, which makes it difficult to model asymmetric relations in taxonomy. In addition, t…

2024

Logit Standardization in Knowledge Distillation

CVPR 2024highlight

Knowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However the assumption of a shared temperature between teacher and student implies a mandatory exact match between their logits in terms of logit range and variance…

2024

Omnidirectional Image Super-resolution via Bi-projection Fusion

AAAI 2024technical

With the rapid development of virtual reality, omnidirectional images (ODIs) have attracted much attention from both the industrial community and academia. However, due to storage and transmission limitations, the resolution of current ODIs is often insufficient to provide an immersive virtual reali…

2024

PAD: Patch-Agnostic Defense against Adversarial Patch Attacks

CVPR 2024poster

Adversarial patch attacks present a significant threat to real-world object detectors due to their practical feasibility. Existing defense methods which rely on attack data or prior knowledge struggle to effectively address a wide range of adversarial patches. In this paper we show two inherent char…

2023

High-Resolution Iterative Feedback Network for Camouflaged Object Detection

AAAI 2023technical

Spotting camouflaged objects that are visually assimilated into the background is tricky for both object detection algorithms and humans who are usually confused or cheated by the perfectly intrinsic similarities between the foreground objects and the background surroundings. To tackle this challeng…

2023

IRNeXt: Rethinking Convolutional Network Design for Image Restoration

ICML 2023poster

We present IRNeXt, a simple yet effective convolutional network architecture for image restoration. Recently, Transformer models have dominated the field of image restoration due to the powerful ability of modeling long-range pixels interactions. In this paper, we excavate the potential of the convo…

2023

Lightweight Image Super-Resolution with Superpixel Token Interaction

ICCV 2023poster

Transformer-based methods have demonstrated impressive results on single-image super-resolution (SISR) task. However, self-attention mechanism is computationally expensive when applied to the entire image. As a result, current approaches divide low-resolution input images into small patches, which a…

Cited by 68PDFcodeScholar
2023

MProto: Multi-Prototype Network with Denoised Optimal Transport for Distantly Supervised Named Entity Recognition

EMNLP 2023long main

Distantly supervised named entity recognition (DS-NER) aims to locate entity mentions and classify their types with only knowledge bases or gazetteers and unlabeled corpus. However, distant annotations are noisy and degrade the performance of NER models. In this paper, we propose a noise-robust prot…

Cited by 0SourcecodeScholar
2023

Robust Single Image Reflection Removal Against Adversarial Attacks

CVPR 2023poster

This paper addresses the problem of robust deep single-image reflection removal (SIRR) against adversarial attacks. Current deep learning based SIRR methods have shown significant performance degradation due to unnoticeable distortions and perturbations on input images. For a comprehensive robustnes…

2023

Selective Frequency Network for Image Restoration

ICLR 2023poster

Image restoration aims to reconstruct the latent sharp image from its corrupted counterpart. Besides dealing with this long-standing task in the spatial domain, a few approaches seek solutions in the frequency domain in consideration of the large discrepancy between spectra of sharp/degraded image p…

Cited by 169SourcePDFScholar
2022

Image Dehazing Transformer With Transmission-Aware 3D Position Embedding

CVPR 2022poster

Despite single image dehazing has been made promising progress with Convolutional Neural Networks (CNNs), the inherent equivariance and locality of convolution still bottleneck dehazing performance. Though Transformer has occupied various computer vision tasks, directly leveraging Transformer for im…

Cited by 412PDFcodeScholar
2022

Self-supervised Learning and Adaptation for Single Image Dehazing

IJCAI 2022poster

Existing deep image dehazing methods usually depend on supervised learning with a large number of hazy-clean image pairs which are expensive or difficult to collect. Moreover, dehazing performance of the learned model may deteriorate significantly when the training hazy-clean image pairs are insuffi…

2021

A Comprehensive Survey on Image Dehazing Based on Deep Learning

IJCAI 2021poster

The presence of haze significantly reduces the quality of images. Researchers have designed a variety of algorithms for image dehazing (ID) to restore the quality of hazy images. However, there are few studies that summarize the deep learning (DL) based dehazing technologies. In this paper, we condu…

Cited by 40SourcePDFScholar
2021

ARVo: Learning All-Range Volumetric Correspondence for Video Deblurring

CVPR 2021poster

Video deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly on homography or optical flows to spatially align neighboring blurry frames. However, such explicit approaches are less…

Cited by 83PDFScholar
2021

Benchmarking Ultra-High-Definition Image Super-Resolution

ICCV 2021poster

Increasingly, modern mobile devices allow capturing images at Ultra-High-Definition (UHD) resolution, which includes 4K and 8K images. However, current single image super-resolution (SISR) methods focus on super-resolving images to ones with resolution up to high definition (HD) and ignore higher-re…

Cited by 38PDFScholar
2021

Clustering-Induced Adaptive Structure Enhancing Network for Incomplete Multi-View Data

IJCAI 2021poster

Incomplete multi-view clustering aims to cluster samples with missing views, which has drawn more and more research interest. Although several methods have been developed for incomplete multi-view clustering, they fail to extract and exploit the comprehensive global and local structure of multi-view…

Cited by 41SourcePDFScholar
2021

DCNAS: Densely Connected Neural Architecture Search for Semantic Image Segmentation

CVPR 2021poster

Existing NAS methods for dense image prediction tasks usually compromise on restricted search space or search on proxy task to meet the achievable computational demands. To allow as wide as possible network architectures and avoid the gap between realistic and proxy setting, we propose a novel Dense…

Cited by 136PDFScholar
2021

Multi-Scale Separable Network for Ultra-High-Definition Video Deblurring

ICCV 2021poster

Although recent research has witnessed a significant progress on the video deblurring task, these methods struggle to reconcile inference efficiency and visual quality simultaneously, especially on ultra-high-definition (UHD) videos (e.g., 4K resolution). To address the problem, we propose a novel d…

Cited by 36PDFcodeScholar
2021

Pyramid Architecture Search for Real-Time Image Deblurring

ICCV 2021poster

Multi-scale and multi-patch deep models have been shown effective in removing blurs of dynamic scenes. However, these methods still have one major obstacle: manually designing a lightweight and high-efficiency network is challenging and time-consuming. To tackle this problem, we propose a novel debl…

Cited by 48PDFScholar
2021

Ultra-High-Definition Image Dehazing via Multi-Guided Bilateral Learning

CVPR 2021poster

During the last couple of years, convolutional neural networks (CNNs) have achieved significant success in the single image dehazing task. Unfortunately, most existing deep dehazing models have high computational complexity, which hinders their application to high-resolution images, especially for U…

Cited by 248PDFcodeScholar
2021

Ultra-High-Definition Image HDR Reconstruction via Collaborative Bilateral Learning

ICCV 2021poster

Existing single image high dynamic range (HDR) reconstruction attempt to expand the range of luminance. They are not effective to generate plausible textures and colors in the reconstructed results, especially for high-density pixels in ultra-high-definition (UHD) images.To address these problems, w…

Cited by 35PDFScholar
2020

Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding

ECCV 2020poster

Rain is a common natural phenomenon. Taking images in the rain however often results in degraded quality of images, thus compromises the performance of many computer vision systems. Most existing de-rain algorithms use only one single input image and aim to recover a clean image. Few work has exploi…

Cited by 54SourcePDFScholar
2020

Face Super-Resolution Guided by 3D Facial Priors

ECCV 2020poster

State-of-the-art face super-resolution methods employ deep convolutional neural networks to learn a mapping between low- and high-resolution facial patterns by exploring local appearance knowledge. However, most of these methods do not well exploit facial structures and identity information, and str…

Cited by 85SourcePDFScholar
2020

Single Image Super-Resolution via a Holistic Attention Network

ECCV 2020poster

Informative features play a crucial role in the single image super-resolution task. Channel attention has been demonstrated to be effective for preserving information-rich features in each layer. However, channel attention treats each convolution layer as a separate process, which is kind of missing…

Cited by 878SourcePDFScholar
2019

Single Image Deraining: A Comprehensive Benchmark Analysis

CVPR 2019poster

We present a comprehensive study and evaluation of existing single image deraining algorithms, using a new large-scale benchmark consisting of both synthetic and real-world rainy images.This dataset highlights diverse data sources and image contents, and is divided into three subsets (rain streak, r…

Cited by 368PDFcodeScholar
2018

Deep Non-Blind Deconvolution via Generalized Low-Rank Approximation

NeurIPS 2018poster

In this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We firs…

Cited by 99SourcePDFScholar
2018

Gated Fusion Network for Single Image Dehazing

CVPR 2018poster

In this paper, we propose an efficient algorithm to directly restore a clear image from a hazy input. The proposed algorithm hinges on an end-to-end trainable neural network that consists of an encoder and a decoder. The encoder is exploited to capture the context of the derived input images, while…

Cited by 1032SourcePDFScholar
2018

Rendering Portraitures from Monocular Camera and Beyond

ECCV 2018poster

Shallow Depth-of-Field (DoF) is a desirable effect in photography which renders artistic photos. Usually, it requires single-lens reflex cameras and certain photography skills to generate such effects. Recently, dual-lens on cellphones is used to estimate scene depth and simulate DoF effects for por…

Cited by 32SourcePDFScholar
2017

Video Deblurring via Semantic Segmentation and Pixel-Wise Non-Linear Kernel

ICCV 2017poster

Video deblurring is a challenging problem as the blur is complex and usually caused by the combination of camera shakes, object motions, and depth variations. Optical flow can be used for kernel estimation since it predicts motion trajectories. However, the estimates are often inaccurate in complex…

Cited by 117PDFScholar