← Search

Lingchen Sun

13 accepted papers

2026

AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation

AAAI 2026technical

Single-image-to-3D models typically follow a sequential generation and reconstruction workflow. However, intermediate multi-view images synthesized by pre-trained generation models often lack cross-view consistency (CVC), significantly degrading 3D reconstruction performance. While recent methods at

Cited by 0SourcePDFScholar
2026

GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution

CVPR 2026

Recently, reinforcement learning (RL) has been employed for improving generative image super-resolution (ISR) performance. However, the current efforts are focused on multi-step generative ISR, while one-step generative ISR remains underexplored due to its limited stochasticity. In addition, RL meth

Cited by 0SourcecodeScholar
2026

Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement

CVPR 2026

Although recent 3D-native generators have made great progress in synthesizing reliable geometry, they still fall short in achieving realistic appearances. A key obstacle lies in the lack of diverse and high-quality real-world 3D assets with rich surface details, since capturing such data is intrinsi

Cited by 0SourcecodeScholar
2026

VOSR: A Vision-Only Generative Model for Image Super-Resolution

CVPR 2026

Large-scale pre-trained text-to-image (T2I) diffusion models, such as Stable Diffusion, can be finetuned for image super-resolution (SR) with highly realistic details. While impressive, pre-training such multi-modal models demands billions of high-quality text-image pairs and substantial computation

Cited by 0SourcecodeScholar
2025

DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution

NeurIPS 2025poster

Benefiting from pre-trained text-to-image (T2I) diffusion models, real-world image super-resolution (Real-ISR) methods can synthesize rich and realistic details. However, due to the inherent stochasticity of T2I models, different noise inputs often lead to outputs with varying perceptual quality. Al…

Cited by 0SourceScholar
2025

Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training

ICCV 2025poster

Impressive results on real-world image super-resolution (Real-ISR) have been achieved by employing pre-trained stable diffusion (SD) models. However, one critical issue of such methods lies in their poor reconstruction of image fine structures, such as small characters and textures, due to the aggre…

2025

GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation

NeurIPS 2025poster

Effective and efficient tokenization plays an important role in image representation and generation. Conventional methods, constrained by uniform 2D/1D grid tokenization, are inflexible to represent regions with varying shapes and textures and at different locations, limiting their efficacy of featu…

Cited by 0SourcecodeScholar
2025

InstructRestore: Region-Customized Image Restoration with Human Instructions

NeurIPS 2025poster

Despite the significant progress in diffusion prior-based image restoration for real-world scenarios, most existing methods apply uniform processing to the entire image, lacking the capability to perform region-customized image restoration according to user preferences. In this work, we propose a ne…

Cited by 0SourcecodeScholar
2025

One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution

NeurIPS 2025poster

It is a challenging problem to reproduce rich spatial details while maintaining temporal consistency in real-world video super-resolution (Real-VSR), especially when we leverage pre-trained generative models such as stable diffusion (SD) for realistic details synthesis. Existing SD-based Real-VSR me…

Cited by 0SourceScholar
2025

Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach

CVPR 2025poster

Diffusion prior-based methods have shown impressive results in real-world image super-resolution (SR). However, most existing methods entangle pixel-level and semantic-level SR objectives in the training process, struggling to balance pixel-wise fidelity and perceptual quality. Meanwhile, users have…

2024

One-Step Effective Diffusion Network for Real-World Image Super-Resolution

NeurIPS 2024poster

The pre-trained text-to-image diffusion models have been increasingly employed to tackle the real-world image super-resolution (Real-ISR) problem due to their powerful generative image priors. Most of the existing methods start from random noise to reconstruct the high-quality (HQ) image under the g…

2024

SeeSR: Towards Semantics-Aware Real-World Image Super-Resolution

CVPR 2024poster

Owe to the powerful generative priors the pre-trained text-to-image (T2I) diffusion models have become increasingly popular in solving the real-world image super-resolution problem. However as a consequence of the heavy quality degradation of input low-resolution (LR) images the destruction of local…

2023

Joint HDR Denoising and Fusion: A Real-World Mobile HDR Image Dataset

CVPR 2023poster

Mobile phones have become a ubiquitous and indispensable photographing device in our daily life, while the small aperture and sensor size make mobile phones more susceptible to noise and over-saturation, resulting in low dynamic range (LDR) and low image quality. It is thus crucial to develop high d…