← Search

Jinshan Pan

69 accepted papers

2026

Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis

ICLR 2026poster

This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive generative augmentation often preserves the biases it aims to solve. Moreover, our analysis reveals that simply generatin…

Cited by 0SourcecodeScholar
2026

Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolution

CVPR 2026

Recent diffusion-based one-step methods have shown remarkable progress in the field of image super-resolution, yet they remain constrained by three critical limitations: (1) inferior fidelity performance caused by the information loss from compression encoding of low-quality (LQ) inputs; (2) insuffi

Cited by 0SourcecodeScholar
2026

FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model

CVPR 2026

Recent studies have witnessed significant advances in image restoration foundation models driven by improvements in the scale and quality of pre-training data. In this work, we find that the data mixture proportions from different restoration tasks are also a critical factor directly determining the

Cited by 0SourcecodeScholar
2026

History to Future: Evolving Agent with Experience and Thought for Zero-shot Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution

Cited by 0SourceScholar
2026

Revisiting Learning with Noisy Labels: Active Forgetting and Noise Suppression

CVPR 2026

Learning with noisy labels (LNL) has received growing attention, with most prior work following the paradigm of clean-sample reliance (e.g., sample selection). However, this reliance also imposes intrinsic limitations, as overfitting to even a few noisy samples is inevitable, creating a major bottle

Cited by 0SourcecodeScholar
2026

STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution

CVPR 2026

We present STCDiT, a video super-resolution framework built upon a pre-trained video diffusion model, aiming to restore structurally faithful and temporally stable videos from degraded inputs, even under complex camera motions. The main challenges lie in maintaining temporal stability during reconst

Cited by 0SourceScholar
2025

CA2C: A Prior-Knowledge-Free Approach for Robust Label Noise Learning via Asymmetric Co-learning and Co-training

ICCV 2025poster

Label noise learning (LNL), a practical challenge in real-world applications, has recently attracted significant attention. While demonstrating promising effectiveness, existing LNL approaches typically rely on various forms of prior knowledge, such as noise rates or thresholds, to sustain performan…

Cited by 0SourcePDFScholar
2025

DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-Resolution

CVPR 2025poster

Recent RGB-guided depth super-resolution methods have achieved impressive performance under the assumption of fixed and known degradation (e.g., bicubic downsampling). However, in real-world scenarios, captured depth data often suffer from unconventional and unknown degradation due to sensor limitat…

Cited by 0SourcePDFScholar
2025

DeblurDiff: Real-Word Image Deblurring with Generative Diffusion Models

NeurIPS 2025poster

Diffusion models have achieved significant progress in image generation and the pre-trained Stable Diffusion (SD) models are helpful for image deblurring by providing clear image priors. However, directly using a blurry image or a pre-deblurred one as a conditional control for SD will either hinder…

Cited by 0SourceScholar
2025

Devil is in the Uniformity: Exploring Diverse Learners within Transformer for Image Restoration

ICCV 2025poster

Transformer-based approaches have gained significant attention in image restoration, where the core component, i.e, Multi-Head Attention (MHA), plays a crucial role in capturing diverse features and recovering high-quality results. In MHA, heads perform attention calculation independently from unifo…

2025

Efficient Concertormer for Image Deblurring and Beyond

ICCV 2025poster

The Transformer architecture has excelled in NLP and vision tasks, but its self-attention complexity grows quadratically with image size, making high-resolution tasks computationally expensive. We introduce Concertormer, featuring Concerto Self-Attention (CSA) for image deblurring. CSA splits self-a…

2025

Efficient Video Super-Resolution for Real-time Rendering with Decoupled G-buffer Guidance

CVPR 2025poster

Latency is a key driver for real-time rendering applications, making super-resolution techniques increasingly popular to accelerate rendering processes. In contrast to existing methods that directly concatenate low-resolution frames and G-buffers as input without discrimination, we develop an asymme…

2025

Efficient Visual State Space Model for Image Deblurring

CVPR 2025poster

Convolutional neural networks (CNNs) and Vision Transformers (ViTs) have achieved excellent performance in image restoration. While ViTs generally outperform CNNs by effectively capturing long-range dependencies and input-specific characteristics, their computational complexity increases quadratical…

2025

FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolution

CVPR 2025poster

Faithful image super-resolution (SR) not only needs to recover images that appear realistic, similar to image generation tasks, but also requires that the restored images maintain fidelity and structural consistency with the input. To this end, we propose a simple and effective method, named FaithDi…

Cited by 4SourcePDFScholar
2025

FlareX: A Physics-Informed Dataset for Lens Flare Removal via 2D Synthesis and 3D Rendering

NeurIPS 2025poster

Lens flare occurs when shooting towards strong light sources, significantly degrading the visual quality of images. Due to the difficulty in capturing flare-corrupted and flare-free image pairs in the real world, existing datasets are typically synthesized in 2D by overlaying artificial flare templa…

Cited by 0SourceScholar
2025

FoundIR: Unleashing Million-scale Training Data to Advance Foundation Models for Image Restoration

ICCV 2025poster

Despite the significant progress made by all-in-one models in universal image restoration, existing methods suffer from a generalization bottleneck in real-world scenarios, as they are mostly trained on small-scale synthetic datasets with limited degradations. Therefore, large-scale high-quality rea…

Cited by 0SourcePDFScholar
2025

Frequency Domain-Based Diffusion Model for Unpaired Image Dehazing

ICCV 2025poster

Unpaired image dehazing has attracted increasing attention due to its flexible data requirements during model training. Dominant methods based on contrastive learning not only introduce haze-unrelated content information, but also ignore haze-specific properties in the frequency domain (i.e., haze-r…

Cited by 0SourcePDFScholar
2025

Intra and Inter Parser-Prompted Transformers for Effective Image Restoration

AAAI 2025technical

We propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Network (IRNet) for restoring images from degraded observations and a Parser-Prompted…

2025

Learning Deblurring Texture Prior from Unpaired Data with Diffusion Model

ICCV 2025poster

Since acquiring large amounts of realistic blurry-sharp image pairs is difficult and expensive, learning blind image deblurring from unpaired data is a more practical and promising solution. Unfortunately, most existing approaches only use adversarial learning to bridge the gap from blurry domains t…

Cited by 0SourcePDFScholar
2025

PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution

ICCV 2025poster

While diffusion models significantly improve the perceptual quality of super-resolved images, they usually require a large number of sampling steps, resulting in high computational costs and long inference times. Recent efforts have explored reasonable acceleration schemes by reducing the number of…

2025

Rethinking Nighttime Image Deraining via Learnable Color Space Transformation

NeurIPS 2025poster

Compared to daytime image deraining, nighttime image deraining poses significant challenges due to inherent complexities of nighttime scenarios and the lack of high-quality datasets that accurately represent the coupling effect between rain and illumination. In this paper, we rethink the task of nig…

Cited by 0SourcecodeScholar
2024

Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image Restoration

CVPR 2024poster

Transformer-based approaches have achieved promising performance in image restoration tasks given their ability to model long-range dependencies which is crucial for recovering clear images. Though diverse efficient attention mechanism designs have addressed the intensive computations associated wit…

2024

Bidirectional Multi-Scale Implicit Neural Representations for Image Deraining

CVPR 2024poster

How to effectively explore multi-scale representations of rain streaks is important for image deraining. In contrast to existing Transformer-based methods that depend mostly on single-scale rain appearance we develop an end-to-end multi-scale Transformer that leverages the potentially useful feature…

2024

Correlation Matching Transformation Transformers for UHD Image Restoration

AAAI 2024technical

This paper proposes UHDformer, a general Transformer for Ultra-High-Definition (UHD) image restoration. UHDformer contains two learning spaces: (a) learning in high-resolution space and (b) learning in low-resolution space. The former learns multi-level high-resolution features and fuses low-high fe…

2024

Seeing the Unseen: A Frequency Prompt Guided Transformer for Image Restoration

ECCV 2024poster

"How to explore useful features from images as prompts to guide the deep image restoration models is an effective way to solve image restoration. In contrast to mining spatial relations within images as prompt, which leads to characteristics of different frequencies being neglected and further remai…

2024

SelfPromer: Self-Prompt Dehazing Transformers with Depth-Consistency

AAAI 2024technical

This work presents an effective depth-consistency Self-Prompt Transformer, terms as SelfPromer, for image dehazing. It is motivated by an observation that the estimated depths of an image with haze residuals and its clear counterpart vary. Enforcing the depth consistency of dehazed images with clear…

2023

DLGSANet: Lightweight Dynamic Local and Global Self-Attention Networks for Image Super-Resolution

ICCV 2023poster

We propose an effective lightweight dynamic local and global self-attention network (DLGSANet) to solve image super-resolution. Our method explores the properties of Transformers while having low computational costs. Motivated by the network designs of Transformers, we develop a simple yet effective…

Cited by 62PDFcodeScholar
2023

Efficient Frequency Domain-Based Transformers for High-Quality Image Deblurring

CVPR 2023highlight

We present an effective and efficient method that explores the properties of Transformers in the frequency domain for high-quality image deblurring. Our method is motivated by the convolution theorem that the correlation or convolution of two signals in the spatial domain is equivalent to an element…

2023

FFHQ-UV: Normalized Facial UV-Texture Dataset for 3D Face Reconstruction

CVPR 2023poster

We present a large-scale facial UV-texture dataset that contains over 50,000 high-quality texture UV-maps with even illuminations, neutral expressions, and cleaned facial regions, which are desired characteristics for rendering realistic 3D face models under different lighting conditions. The datase…

2023

Hybrid CNN-Transformer Feature Fusion for Single Image Deraining

AAAI 2023technical

Since rain streaks exhibit diverse geometric appearances and irregular overlapped phenomena, these complex characteristics challenge the design of an effective single image deraining model. To this end, rich local-global information representations are increasingly indispensable for better satisfyin…

2023

Learning a Sparse Transformer Network for Effective Image Deraining

CVPR 2023highlight

Transformers-based methods have achieved significant performance in image deraining as they can model the non-local information which is vital for high-quality image reconstruction. In this paper, we find that most existing Transformers usually use all similarities of the tokens from the query-key p…

2023

PromptRestorer: A Prompting Image Restoration Method with Degradation Perception

NeurIPS 2023poster

We show that raw degradation features can effectively guide deep restoration models, providing accurate degradation priors to facilitate better restoration. While networks that do not consider them for restoration forget gradually degradation during the learning process, model capacity is severely h…

Cited by 60SourcePDFScholar
2023

SVMV: Spatiotemporal Variance-Supervised Motion Volume for Video Frame Interpolation

ICASSP 2023accepted

High-performance video frame interpolation is challenging for complex scenes with diverse motion and occlusion characteristics. Existing methods, deploying off-the-shelf flow estimators to acquire initial characterizations refined by multiple subsequent models, often require heavy network architectu…

Cited by 0SourceScholar
2023

Spatially-Adaptive Feature Modulation for Efficient Image Super-Resolution

ICCV 2023poster

Although deep learning-based solutions have achieved impressive reconstruction performance in image super-resolution (SR), these models are generally large, with complex architectures, making them incompatible with low-power devices with many computational and memory constraints. To overcome these c…

Cited by 142PDFcodeScholar
2022

Deep Recurrent Neural Network with Multi-Scale Bi-directional Propagation for Video Deblurring

AAAI 2022technical

The success of the state-of-the-art video deblurring methods stems mainly from implicit or explicit estimation of alignment among the adjacent frames for latent video restoration. However, due to the influence of the blur effect, estimating the alignment information from the blurry adjacent frames i…

2022

Learning Discriminative Shrinkage Deep Networks for Image Deconvolution

ECCV 2022poster

"Most existing methods usually formulate the non-blind deconvolution problem into a maximum-a-posteriori framework and address it by manually designing a variety of regularization terms and data terms of the latent clear images. However, explicitly designing these two terms is quite challenging and…

2022

Online-Updated High-Order Collaborative Networks for Single Image Deraining

AAAI 2022technical

Single image deraining is an important and challenging task for some downstream artificial intelligence applications such as video surveillance and self-driving systems. Most of the existing deep-learning-based methods constrain the network to generate derained images but few of them explore feature…

Cited by 26SourcePDFScholar
2022

Unpaired Deep Image Deraining Using Dual Contrastive Learning

CVPR 2022poster

Learning single image deraining (SID) networks from an unpaired set of clean and rainy images is practical and valuable as acquiring paired real-world data is almost infeasible. However, without the paired data as the supervision, learning a SID network is challenging. Moreover, simply using existin…

Cited by 196PDFScholar
2021

Learning To Restore Hazy Video: A New Real-World Dataset and a New Method

CVPR 2021poster

Most of the existing deep learning-based dehazing methods are trained and evaluated on the image dehazing datasets, where the dehazed images are generated by only exploiting the information from the corresponding hazy ones. On the other hand, the video dehazing algorithms, which can acquire more sat…

Cited by 103PDFScholar
2021

Learning a Non-Blind Deblurring Network for Night Blurry Images

CVPR 2021poster

Deblurring night blurry images is difficult, because the common-used blur model based on the linear convolution operation does not hold in this situation due to the influence of saturated pixels. In this paper, we propose a non-blind deblurring network (NBDN) to restore night blurry images. To mitig…

Cited by 36PDFScholar
2020

Learning Event-Driven Video Deblurring and Interpolation

ECCV 2020poster

Event-based sensors, which have a response if the change of pixel intensity exceeds a triggering threshold, can capture high-speed motion with microsecond accuracy. Assisted by an event camera, we can generate high frame-rate sharp videos from low frame-rate blurry ones captured by an intensity came…

Cited by 154SourcePDFScholar
2020

Multi-Scale Boosted Dehazing Network With Dense Feature Fusion

CVPR 2020poster

In this paper, we propose a Multi-Scale Boosted Dehazing Network with Dense Feature Fusion based on the U-Net architecture. The proposed method is designed based on two principles, boosting and error feedback, and we show that they are suitable for the dehazing problem. By incorporating the Strength…

Cited by 1034PDFcodeScholar
2019

DAVANet: Stereo Deblurring With View Aggregation

CVPR 2019oral

Nowadays stereo cameras are more commonly adopted in emerging devices such as dual-lens smartphones and unmanned aerial vehicles. However, they also suffer from blurry images in dynamic scenes which leads to visual discomfort and hampers further image processing. Previous works have succeeded in mon…

Cited by 119PDFScholar
2019

Spatially Variant Linear Representation Models for Joint Filtering

CVPR 2019poster

Joint filtering mainly uses an additional guidance image as a prior and transfers its structures to the target image in the filtering process. Different from existing algorithms that rely on locally linear models or hand-designed objective functions to extract the structural information from the gui…

Cited by 53PDFScholar
2019

Spatio-Temporal Filter Adaptive Network for Video Deblurring

ICCV 2019poster

Video deblurring is a challenging task due to the spatially variant blur caused by camera shake, object motions, and depth variations, etc. Existing methods usually estimate optical flow in the blurry video to align consecutive frames or approximate blur kernels. However, they tend to generate artif…

Cited by 247PDFScholar
2018

Deep Non-Blind Deconvolution via Generalized Low-Rank Approximation

NeurIPS 2018poster

In this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We firs…

Cited by 99SourcePDFScholar
2018

Dynamic Scene Deblurring Using Spatially Variant Recurrent Neural Networks

CVPR 2018poster

Due to the spatially variant blur caused by camera shake and object motions under different scene depths, deblurring images captured from dynamic scenes is challenging. Although recent works based on deep neural networks have shown great progress on this problem, their models are usually large and c…

Cited by 466SourcePDFScholar
2018

Gated Fusion Network for Single Image Dehazing

CVPR 2018poster

In this paper, we propose an efficient algorithm to directly restore a clear image from a hazy input. The proposed algorithm hinges on an end-to-end trainable neural network that consists of an encoder and a decoder. The encoder is exploited to capture the context of the derived input images, while…

Cited by 1032SourcePDFScholar
2018

Learning Dual Convolutional Neural Networks for Low-Level Vision

CVPR 2018poster

In this paper, we propose a general dual convolutional neural network (DualCNN) for low-level vision problems, e.g., super-resolution, edge-preserving filtering, deraining and dehazing. These problems usually involve the estimation of two components of the target signals: structures and details. Mot…

Cited by 230SourcePDFScholar
2018

Learning a Discriminative Prior for Blind Image Deblurring

CVPR 2018poster

We present an effective blind image deblurring method based on a data-driven discriminative prior. Our work is motivated by the fact that a good image prior should favor clear images over blurred images. To obtain such an image prior for deblurring, we formulate the image prior as a binary classifie…

Cited by 195SourcePDFScholar
2018

Single Image Dehazing via Conditional Generative Adversarial Network

CVPR 2018poster

In this paper, we present an algorithm to directly restore a clear image from a hazy image. This problem is highly ill-posed and most existing algorithms often use hand-crafted features, e.g., dark channel, color disparity, maximum contrast, to estimate transmission maps and then atmospheric lights.…

Cited by 511SourcePDFScholar
2017

Learning Discriminative Data Fitting Functions for Blind Image Deblurring

ICCV 2017poster

Solving blind image deblurring usually requires defining a data fitting function and image priors. While existing algorithms mainly focus on developing image priors for blur kernel estimation and non-blind deconvolution, only a few methods consider the effect of data fitting functions. In contrast t…

Cited by 38PDFScholar
2017

Learning Fully Convolutional Networks for Iterative Non-Blind Deconvolution

CVPR 2017poster

In this paper, we propose a fully convolutional network for iterative non-blind deconvolution. We decompose the non-blind deconvolution problem into image denoising and image deconvolution. We train a FCNN to remove noise in the gradient domain and use the learned gradients to guide the image deconv…

Cited by 215PDFScholar
2017

Learning to Super-Resolve Blurry Face and Text Images

ICCV 2017poster

We present an algorithm to directly restore a clear high-resolution image from a blurry low-resolution input. This problem is highly ill-posed and the basic assumptions for existing super-resolution methods (requiring clear input) and deblurring methods (requiring high-resolution input) no longer ho…

Cited by 279PDFScholar
2017

Video Deblurring via Semantic Segmentation and Pixel-Wise Non-Linear Kernel

ICCV 2017poster

Video deblurring is a challenging problem as the blur is complex and usually caused by the combination of camera shakes, object motions, and depth variations. Optical flow can be used for kernel estimation since it predicts motion trajectories. However, the estimates are often inaccurate in complex…

Cited by 117PDFScholar
2016

Robust Kernel Estimation With Outliers Handling for Image Deblurring

CVPR 2016poster

Estimating blur kernels from real world images is a challenging problem as the linear image formation assumption does not hold when significant outliers, such as saturated pixels and non-Gaussian noise, are present. While some existing non-blind deblurring algorithms can deal with outliers to a cert…

Cited by 130PDFScholar