← Search

Chongyi Li

57 accepted papers

2026

DNF-SR: Dual-Input and Negative-Aware Feature Fine-Tuning for Real-World Image Super-Resolution

CVPR 2026

Benefiting from the powerful generative priors of diffusion models, diffusion-based real-world image super-resolution (Real-ISR) methods have demonstrated impressive performance.To achieve efficient Real-ISR, several recent works have designed one-step diffusion-based models.Howerver, unmediatedly f

Cited by 0SourcecodeScholar
2026

EvalMuse-40K: A Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Alignment Evaluation

AAAI 2026technical

Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated methods emerge to evaluate the image-text alignment capabilities of generative models. However, the performance comparison among these automated methods is constrained by the limited scale o

Cited by 0SourcePDFScholar
2026

Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory

ICML 2026poster

We propose **Infinite-World**, a robust interactive world model capable of maintaining coherent visual memory over **1000+ frames** in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with perfect ground-truth, they lack an effective trainin…

Cited by 0SourceScholar
2026

Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation

CVPR 2026

Generating high-fidelity human videos that match user-specified identities is important yet challenging in the field of generative AI.Existing methods often rely on an excessive number of training parameters and lack compatibility with other AIGC tools.In this paper, we propose Stand-In, a lightweig

Cited by 0SourcecodeScholar
2026

Time-Aware One Step Diffusion Network for Real-World Image Super-Resolution

CVPR 2026

Diffusion-based real-world image super-resolution (Real-ISR) methods have demonstrated impressive performance. To achieve efficient Real-ISR, many works employ Variational Score Distillation (VSD) to distill a pre-trained stable-diffusion (SD) model for one-step SR with a fixed timestep. However, si

Cited by 0SourcecodeScholar
2026

VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame Interpolation

AAAI 2026technical

Due to large pixel movement and high computational cost, estimating the motion of high-resolution frames is challenging. Thus, most flow-based Video Frame Interpolation (VFI) methods first predict bidirectional flows at low resolution and then use high-magnification upsampling (e.g., bilinear) to ob

Cited by 0SourcePDFScholar
2026

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal

CVPR 2026

Recent advances in Diffusion Transformer (DiT)-based video generation technologies have shown impressive results for video object removal. However, these methods still suffer from substantial inference latency. For instance, although MiniMax Remover achieves state-of-the-art visual quality, it opera

Cited by 0SourcecodeScholar
2025

A Diffusion-Based Framework for Occluded Object Movement

AAAI 2025technical

Seamlessly moving objects within a scene is a common requirement for image editing, but it is still a challenge for existing editing methods. Especially for real-world images, the occlusion situation further increases the difficulty. The main difficulty is that the occluded portion needs to be compl…

Cited by 0SourcePDFScholar
2025

Classic Video Denoising in a Machine Learning World: Robust, Fast, and Controllable

CVPR 2025poster

Denoising is a crucial step in many video processing pipelines such as in interactive editing, where high quality, speed, and user control are essential. While recent approaches achieve significant improvements in denoising quality by leveraging deep learning, they are prone to unexpected failures d…

Cited by 0SourcePDFScholar
2025

DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data

NeurIPS 2025poster

We present **DIPO**, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state. Compared to the single-image approach, our dual-image input imposes only a modest overhead for da…

Cited by 0SourcecodeScholar
2025

DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution

ICCV 2025poster

Large-scale pre-trained diffusion models are becoming increasingly popular in solving the Real-World Image Super-Resolution (Real-ISR) problem because of their rich generative priors. The recent development of diffusion transformer (DiT) has witnessed overwhelming performance over the traditional UN…

Cited by 0SourcePDFScholar
2025

DiffRetouch: Using Diffusion to Retouch on the Shoulder of Experts

AAAI 2025technical

Image retouching aims to enhance the visual quality of photos. Considering the different aesthetic preferences of users, the target of retouching is subjective. However, current retouching methods mostly adopt deterministic models, which not only neglects the style diversity in the expert-retouched…

Cited by 0SourcePDFScholar
2025

FaceMe: Robust Blind Face Restoration with Personal Identification

AAAI 2025technical

Blind face restoration is a highly ill-posed problem due to the lack of necessary context. Although existing methods produce high-quality outputs, they often fail to faithfully preserve the individual's identity. In this paper, we propose a personalized face restoration method, FaceMe, based on a di…

2025

Iterative Predictor-Critic Code Decoding for Real-World Image Dehazing

CVPR 2025poster

We propose a novel Iterative Predictor-Critic Code Decoding framework for real-world image dehazing, abbreviated as IPC-Dehaze, which leverages the high-quality codebook prior encapsulated in a pre-trained VQGAN. Apart from previous codebook-based methods that rely on one-shot decoding, our method u…

2025

Joint Semantic and Rendering Enhancements in 3D Gaussian Modeling with Anisotropic Local Encoding

ICCV 2025poster

Recent works propose extending 3DGS with semantic feature vectors for simultaneous semantic segmentation and image rendering. However, these methods often treat the semantic and rendering branches separately, relying solely on 2D supervision while ignoring the 3D Gaussian geometry. Moreover, current…

2025

MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data Generation

ICCV 2025poster

Recent advancements in Face Image Quality Assessment (FIQA) models trained on real large-scale face datasets are pivotal in guaranteeing precise face recognition in unrestricted scenarios. Regrettably, privacy concerns lead to the discontinuation of real datasets, underscoring the pressing need for…

2025

UltraLED: Learning to See Everything in Ultra-High Dynamic Range Scenes

NeurIPS 2025poster

Ultra-high dynamic range (UHDR) scenes exhibit pronounced exposure disparities between bright and dark regions. Such conditions are Ultra-high dynamic range (UHDR) scenes exhibit significant exposure disparities between bright and dark regions. Such conditions are commonly encountered in nighttime s…

Cited by 0SourcecodeScholar
2024

AMSP-UOD: When Vortex Convolution and Stochastic Perturbation Meet Underwater Object Detection

AAAI 2024technical

In this paper, we present a novel Amplitude-Modulated Stochastic Perturbation and Vortex Convolutional Network, AMSP-UOD, designed for underwater object detection. AMSP-UOD specifically addresses the impact of non-ideal imaging factors on detection accuracy in complex underwater environments. To mit…

2024

Adaptive Window Pruning for Efficient Local Motion Deblurring

ICLR 2024poster

Local motion blur commonly occurs in real-world photography due to the mixing between moving objects and stationary backgrounds during exposure. Existing image deblurring methods predominantly focus on global deblurring, inadvertently affecting the sharpness of backgrounds in locally blurred images…

Cited by 7SourcePDFScholar
2024

CLIB-FIQA: Face Image Quality Assessment with Confidence Calibration

CVPR 2024poster

Face Image Quality Assessment (FIQA) is pivotal for guaranteeing the accuracy of face recognition in unconstrained environments. Recent progress in deep quality-fitting-based methods that train models to align with quality anchors has shown promise in FIQA. However these methods heavily depend on a…

2024

Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring

CVPR 2024poster

Existing joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models they typically rely on the assumed degra…

2024

Kalman-Inspired Feature Propagation for Video Face Super-Resolution

ECCV 2024poster

"Despite the promising progress of face image super-resolution, video face super-resolution remains relatively under-explored. Existing approaches either adapt general video super-resolution networks to face datasets or apply established face image super-resolution models independently on individual…

2024

LAMP: Learn A Motion Pattern for Few-Shot Video Generation

CVPR 2024poster

In this paper we present a few-shot text-to-video framework LAMP which enables a text-to-image diffusion model to Learn A specific Motion Pattern with 8 16 videos on a single GPU. Unlike existing methods which require a large number of training resources or learn motions that are precisely aligned w…

2024

Learning Inclusion Matching for Animation Paint Bucket Colorization

CVPR 2024poster

Colorizing line art is a pivotal task in the production of hand-drawn cel animation. This typically involves digital painters using a paint bucket tool to manually color each segment enclosed by lines based on RGB values predetermined by a color designer. This frame-by-frame process is both arduous…

2024

Lighting Every Darkness with 3DGS: Fast Training and Real-Time Rendering for HDR View Synthesis

NeurIPS 2024poster

Volumetric rendering-based methods, like NeRF, excel in HDR view synthesis from RAW images, especially for nighttime scenes. They suffer from long training times and cannot perform real-time rendering due to dense sampling requirements. The advent of 3D Gaussian Splatting (3DGS) enables real-time re…

2024

Synergistic Multiscale Detail Refinement via Intrinsic Supervision for Underwater Image Enhancement

AAAI 2024technical

Visually restoring underwater scenes primarily involves mitigating interference from underwater media. Existing methods ignore the inherent scale-related characteristics in underwater scenes. Therefore, we present the synergistic multi-scale detail refinement via intrinsic supervision (SMDR-IS) for…

2023

DNF: Decouple and Feedback Network for Seeing in the Dark

CVPR 2023highlight

The exclusive properties of RAW data have shown great potential for low-light image enhancement. Nevertheless, the performance is bottlenecked by the inherent limitations of existing architectures in both single-stage and multi-stage methods. Mixed mapping across two different domains, noise-to-clea…

2023

Embedding Fourier for Ultra-High-Definition Low-Light Image Enhancement

ICLR 2023top-5%

Ultra-High-Definition (UHD) photo has gradually become the standard configuration in advanced imaging devices. The new standard unveils many issues in existing approaches for low-light image enhancement (LLIE), especially in dealing with the intricate issue of joint luminance enhancement and noise r…

2023

Empowering Low-Light Image Enhancer through Customized Learnable Priors

ICCV 2023poster

Deep neural networks have achieved remarkable progress in enhancing low-light images by improving their brightness and eliminating noise. However, most existing methods construct end-to-end mapping networks heuristically, neglecting the intrinsic prior of image enhancement task and lacking transpare…

Cited by 46PDFcodeScholar
2023

Exploring Temporal Frequency Spectrum in Deep Video Deblurring

ICCV 2023poster

Video deblurring aims to restore the latent video frames from their blurred counterparts. Despite the remarkable progress, most promising video deblurring methods only investigate the temporal priors in the spatial domain and rarely explore their its potential in the frequency domain. In this paper,…

Cited by 24PDFScholar
2023

FouriDown: Factoring Down-Sampling into Shuffling and Superposing

NeurIPS 2023poster

Spatial down-sampling techniques, such as strided convolution, Gaussian, and Nearest down-sampling, are essential in deep neural networks. In this study, we revisit the working mechanism of the spatial down-sampling family and analyze the biased effects caused by the static weighting strategy employ…

2023

Generating Aligned Pseudo-Supervision From Non-Aligned Data for Image Restoration in Under-Display Camera

CVPR 2023poster

Due to the difficulty in collecting large-scale and perfectly aligned paired training data for Under-Display Camera (UDC) image restoration, previous methods resort to monitor-based image systems or simulation-based methods, sacrificing the realness of the data and introducing domain gaps. In this w…

2023

Improving Lens Flare Removal with General-Purpose Pipeline and Multiple Light Sources Recovery

ICCV 2023poster

When taking images against strong light sources, the resulting images often contain heterogeneous flare artifacts. These artifacts can importantly affect image visual quality and downstream computer vision tasks. While collecting real data pairs of flare-corrupted/flare-free images for training flar…

Cited by 26PDFcodeScholar
2023

Iterative Prompt Learning for Unsupervised Backlit Image Enhancement

ICCV 2023oral

We propose a novel unsupervised backlit image enhancement method, abbreviated as CLIP-LIT, by exploring the potential of Contrastive Language-Image Pre-Training (CLIP) for pixel-level image enhancement. We show that the open-world CLIP prior not only aids in distinguishing between backlit and well-l…

Cited by 140PDFScholar
2023

Learned Image Reasoning Prior Penetrates Deep Unfolding Network for Panchromatic and Multi-spectral Image Fusion

ICCV 2023poster

The success of deep neural networks for pan-sharpening is commonly in a form of black box, lacking transparency and interpretability. To alleviate this issue, we propose a novel model-driven deep unfolding framework with image reasoning prior tailored for the pan-sharpening task. Different from exis…

Cited by 10PDFScholar
2023

Learning Semantic-Aware Knowledge Guidance for Low-Light Image Enhancement

CVPR 2023poster

Low-light image enhancement (LLIE) investigates how to improve illumination and produce normal-light images. The majority of existing methods improve low-light images via a global and uniform manner, without taking into account the semantic information of different regions. Without semantic priors,…

2023

Lighting Every Darkness in Two Pairs: A Calibration-Free Pipeline for RAW Denoising

ICCV 2023poster

Calibration-based methods have dominated RAW image denoising under extremely low-light environments. However, these methods suffer from several main deficiencies: 1) the calibration procedure is laborious and time-consuming, 2) denoisers for different cameras are difficult to transfer, and 3) the di…

Cited by 23PDFScholar
2023

Nighttime Smartphone Reflective Flare Removal Using Optical Center Symmetry Prior

CVPR 2023highlight

Reflective flare is a phenomenon that occurs when light reflects inside lenses, causing bright spots or a "ghosting effect" in photos, which can impact their quality. Eliminating reflective flare is highly desirable but challenging. Many existing methods rely on manually designed features to detect…

2023

ProPainter: Improving Propagation and Transformer for Video Inpainting

ICCV 2023poster

Flow-based propagation and spatiotemporal Transformer are two mainstream mechanisms in video inpainting (VI). Despite the effectiveness of these components, they still suffer from some limitations that affect their performance. Previous propagation-based approaches are performed separately either in…

Cited by 104PDFcodeScholar
2023

RIDCP: Revitalizing Real Image Dehazing via High-Quality Codebook Priors

CVPR 2023poster

Existing dehazing approaches struggle to process real-world hazy images owing to the lack of paired real data and robust priors. In this work, we present a new paradigm for real image dehazing from the perspectives of synthesizing more realistic hazy data and introducing more robust priors into the…

2023

Training Your Image Restoration Network Better with Random Weight Network as Optimization Function

NeurIPS 2023poster

The blooming progress made in deep learning-based image restoration has been largely attributed to the availability of high-quality, large-scale datasets and advanced network structures. However, optimization functions such as L_1 and L_2 are still de facto. In this study, we propose to investigate…

Cited by 1SourcePDFScholar
2023

Transition-constant Normalization for Image Enhancement

NeurIPS 2023spotlight

Normalization techniques that capture image style by statistical representation have become a popular component in deep neural networks. Although image enhancement can be considered as a form of style transformation, there has been little exploration of how normalization affect the enhancement perfo…

2023

Troubleshooting Ethnic Quality Bias with Curriculum Domain Adaptation for Face Image Quality Assessment

ICCV 2023poster

Face Image Quality Assessment (FIQA) lays the foundation for ensuring the stability and accuracy of face recognition systems. However, existing FIQA methods mainly formulate quality relationships within the training set to yield quality scores, ignoring the generalization problem caused by ethnic qu…

Cited by 9PDFcodeScholar
2023

Underwater Ranker: Learn Which Is Better and How to Be Better

AAAI 2023technical

In this paper, we present a ranking-based underwater image quality assessment (UIQA) method, abbreviated as URanker. The URanker is built on the efficient conv-attentional image Transformer. In terms of underwater images, we specially devise (1) the histogram prior that embeds the color distribution…

2022

Flare7K: A Phenomenological Nighttime Flare Removal Dataset

NeurIPS 2022accept

Artificial lights commonly leave strong lens flare artifacts on images captured at night. Nighttime flare not only affects the visual quality but also degrades the performance of vision algorithms. Existing flare removal methods mainly focus on removing daytime flares and fail in nighttime. Nighttim…

2022

Image Dehazing Transformer With Transmission-Aware 3D Position Embedding

CVPR 2022poster

Despite single image dehazing has been made promising progress with Convolutional Neural Networks (CNNs), the inherent equivariance and locality of convolution still bottleneck dehazing performance. Though Transformer has occupied various computer vision tasks, directly leveraging Transformer for im…

Cited by 412PDFcodeScholar
2022

LEDNet: Joint Low-Light Enhancement and Deblurring in the Dark

ECCV 2022poster

"Night photography typically suffers from both low light and blurring issues due to the dim environment and the common use of long exposure. While existing light enhancement and deblurring methods could deal with each problem individually, a cascade of such methods cannot work harmoniously to cope w…

2022

Panchromatic and Multispectral Image Fusion via Alternating Reverse Filtering Network

NeurIPS 2022accept

Panchromatic (PAN) and multi-spectral (MS) image fusion, named Pan-sharpening, refers to super-resolve the low-resolution (LR) multi-spectral (MS) images in the spatial domain to generate the expected high-resolution (HR) MS images, conditioning on the corresponding high-resolution PAN images. In th…

Cited by 21SourcePDFScholar
2022

Towards Robust Blind Face Restoration with Codebook Lookup Transformer

NeurIPS 2022accept

Blind face restoration is a highly ill-posed problem that often requires auxiliary guidance to 1) improve the mapping from degraded inputs to desired outputs, or 2) complement high-quality details lost in the inputs. In this paper, we demonstrate that a learned discrete codebook prior in a small pro…

2021

Removing Diffraction Image Artifacts in Under-Display Camera via Dynamic Skip Connection Network

CVPR 2021poster

Recent development of Under-Display Camera (UDC) systems provides a true bezel-less and notch-free viewing experience on smartphones (and TV, laptops, tablets), while allowing images to be captured from the selfie camera embedded underneath. In a typical UDC system, the microstructure of the semi-tr…

Cited by 74PDFcodeScholar
2020

CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object Detection

NeurIPS 2020poster

Co-Salient Object Detection (CoSOD) aims at discovering salient objects that repeatedly appear in a given query group containing two or more relevant images. One challenging issue is how to effectively capture co-saliency cues by modeling and exploiting inter-image relationships. In this paper, we p…

2020

RGB-D Salient Object Detection with Cross-Modality Modulation and Selection

ECCV 2020poster

We present an effective method to progressively integrate and refine the cross-modality complementarities for RGB-D salient object detection (SOD). The proposed network mainly solves two challenging issues: 1) how to effectively integrate the complementary information from RGB image and its correspo…

Cited by 168SourcePDFScholar
2020

Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement

CVPR 2020poster

The paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep network. Our method trains a lightweight deep network, DCE-Net, to estimate pixel-wise and high-order curves for dynamic ran…

Cited by 2093PDFcodeScholar
2016

Single underwater image restoration by blue-green channels dehazing and red channel correction

ICASSP 2016accepted

Restoring underwater image from a single image is know to be ill-posed, and some assumptions made in previous methods are not suitable for many situations. In this paper, we propose a method based on blue-green channels dehazing and red channel correction for underwater image restoration. Firstly, b…

Cited by 0SourceScholar