← Search

Qingnan Fan

29 accepted papers

2026

GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors

ICML 2026poster

Reconstructing 3D scenes using 3D Gaussian Splatting (3DGS) from sparse views is an ill-posed problem due to insufficient information, often resulting in noticeable artifacts. While recent approaches have sought to leverage generative priors to complete information for under-constrained regions, the…

Cited by 0SourceScholar
2026

LiveMoments: Reselected Key Photo Restoration in Live Photos via Reference-guided Diffusion

ICLR 2026poster

Live Photo captures both a high-quality key photo and a short video clip to preserve the precious dynamics around the captured moment. While users may choose alternative frames as the key photo to capture better expressions or timing, these frames often exhibit noticeable quality degradation, as th…

Cited by 0SourcecodeScholar
2026

One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step Diffusion

AAAI 2026technical

We present a novel framework for high-fidelity novel view synthesis (NVS) from sparse images, addressing key limitations in recent feed-forward 3D Gaussian Splatting (3DGS) methods built on Vision Transformer (ViT) backbones. While ViT-based pipelines offer strong geometric priors, they are often co

Cited by 0SourcePDFScholar
2026

Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance

CVPR 2026

Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce **TIGER** (**T**ext-**I**mage **G**uided sup**E**r-**R**esolution), a novel two-stage frame

Cited by 0SourceScholar
2025

BokehDiff: Neural Lens Blur with One-Step Diffusion

ICCV 2025poster

We introduce Bokehdiff, a novel lens blur rendering method that achieves physically accurate and visually appealing outcomes, with the help of generative diffusion prior. Previous methods are bounded by the accuracy of depth estimation, generating artifacts in depth discontinuities. Our method emplo…

2025

CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models

ICCV 2025poster

Text-to-image (T2I) diffusion models excel at generating photorealistic images, but commonly struggle to render accurate spatial relationships described in text prompts. We identify two core issues underlying this common failure: 1) the ambiguous nature of spatial-related data in existing datasets,…

2025

RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution

AAAI 2025technical

Benefiting from their powerful generative capabilities, pretrained diffusion models have garnered significant attention for real-world image super-resolution (Real-SR). Existing diffusion-based SR approaches typically utilize semantic information from degraded images and restoration prompts to activ…

2025

Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization

CVPR 2025poster

Visual localization aims to determine the camera pose of a query image relative to a database of posed images. In recent years, deep neural networks that directly regress camera poses have gained popularity due to their fast inference capabilities. However, existing methods struggle to either genera…

2025

SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos

CVPR 2025highlight

In this paper, we introduce SLAM3R, a novel and effective system for real-time, high-quality, dense 3D reconstruction using RGB videos. SLAM3R provides an end-to-end solution by seamlessly integrating local 3D reconstruction and global coordinate registration through feed-forward neural networks. Gi…

2025

TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-Resolution

CVPR 2025poster

Pre-trained text-to-image diffusion models are increasingly applied to real-world image super-resolution (Real-ISR) task. Given the iterative refinement nature of diffusion models, most existing approaches are computationally expensive. While methods such as SinSR and OSEDiff have emerged to condens…

2025

Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders

NeurIPS 2025poster

The introduction of generative models has significantly advanced image super-resolution (SR) in handling real-world degradations. However, they often incur fidelity-related issues, particularly distorting textual structures. In this paper, we introduce a novel diffusion-based SR framework, namely T…

Cited by 0SourcecodeScholar
2025

Textualize Visual Prompt for Image Editing via Diffusion Bridge

AAAI 2025technical

Visual prompt, a pair of before-and-after edited images, can convey indescribable imagery transformations and prosper in image editing. However, current visual prompt methods rely on a pretrained text-guided image-to-image generative model that requires a triplet of text, before, and after images fo…

Cited by 0SourcePDFScholar
2024

FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion Models

ECCV 2024poster

"Precise image editing with text-to-image models has attracted increasing interest due to their remarkable generative capabilities and user-friendly nature. However, such attempts face the pivotal challenge of misalignment between the intended precise editing target regions and the broader area impa…

2023

3D-Aware Object Goal Navigation via Simultaneous Exploration and Identification

CVPR 2023poster

Object goal navigation (ObjectNav) in unseen environments is a fundamental task for Embodied AI. Agents in existing works learn ObjectNav policies based on 2D maps, scene graphs, or image sequences. Considering this task happens in 3D space, a 3D-aware agent can advance its ObjectNav capability via…

Cited by 47SourcePDFScholar
2023

DualAfford: Learning Collaborative Visual Affordance for Dual-gripper Manipulation

ICLR 2023poster

It is essential yet challenging for future home-assistant robots to understand and manipulate diverse 3D objects in daily human environments. Towards building scalable systems that can perform diverse manipulation tasks over various 3D shapes, recent works have advocated and demonstrated promising r…

Cited by 16SourcePDFScholar
2022

ADeLA: Automatic Dense Labeling With Attention for Viewpoint Shift in Semantic Segmentation

CVPR 2022oral

We describe a method to deal with performance drop in semantic segmentation caused by viewpoint changes within multi-camera systems, where temporally paired images are readily available, but the annotations may only be abundant for a few typical views. Existing methods alleviate performance drop via…

Cited by 6PDFScholar
2022

AdaAfford: Learning to Adapt Manipulation Affordance for 3D Articulated Objects via Few-Shot Interactions

ECCV 2022poster

"Perceiving and interacting with 3D articulated objects, such as cabinets, doors, and faucets, pose particular challenges for future home-assistant robots performing daily tasks in human environments. Besides parsing the articulated parts and joint parameters, researchers recently advocate learning…

Cited by 68SourcePDFScholar
2022

Multi-Robot Active Mapping via Neural Bipartite Graph Matching

CVPR 2022poster

We study the problem of multi-robot active mapping, which aims for complete scene map construction in minimum time steps. The key to this problem lies in the goal position estimation to enable more efficient robot movements. Previous approaches either choose the frontier as the goal position via a m…

Cited by 34PDFScholar
2022

Towards Accurate Active Camera Localization

ECCV 2022poster

"In this work, we tackle the problem of active camera localization, which controls the camera movements actively to achieve an accurate camera pose. The past solutions are mostly based on Markov Localization, which reduces the position-wise camera uncertainty for localization. These approaches local…

2022

VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated Objects

ICLR 2022poster

Perceiving and manipulating 3D articulated objects (e.g., cabinets, doors) in human environments is an important yet challenging task for future home-assistant robots. The space of 3D articulated objects is exceptionally rich in their myriad semantic categories, diverse shape geometry, and complicat…

Cited by 104SourcePDFScholar
2021

CAPTRA: CAtegory-Level Pose Tracking for Rigid and Articulated Objects From Point Clouds

ICCV 2021poster

In this work, we tackle the problem of category-level online pose tracking for objects from point cloud sequences. For the first time, we propose a unified framework that can handle 9DoF object pose tracking for novel rigid object instances as well as per-part pose tracking for articulated objects f…

Cited by 114PDFcodeScholar
2021

Contrastive Multimodal Fusion With TupleInfoNCE

ICCV 2021poster

This paper proposes a method for representation learning of multimodal data using contrastive losses. A traditional approach is to contrast different modalities to learn the information shared between them. However, that approach could fail to learn the complementary synergies between modalities tha…

Cited by 83PDFcodeScholar
2021

Generating Manga From Illustrations via Mimicking Manga Creation Workflow

CVPR 2021poster

We present a framework to generate manga from digital illustrations. In professional mange studios, the manga create workflow consists of three key steps: (1) Artists use line drawings to delineate the structural outlines in manga storyboards. (2) Artists apply several types of regular screentones t…

Cited by 21PDFScholar
2021

Robust Neural Routing Through Space Partitions for Camera Relocalization in Dynamic Indoor Environments

CVPR 2021poster

Localizing the camera in a known indoor environment is a key building block for scene mapping, robot navigation, AR, etc. Recent advances estimate the camera pose via optimization over the 2D/3D-3D correspondences established between the coordinates in 2D/3D camera space and 3D world space. Such a m…

Cited by 32PDFcodeScholar
2020

Generative 3D Part Assembly via Dynamic Graph Learning

NeurIPS 2020poster

Autonomous part assembly is a challenging yet crucial task in 3D computer vision and robotics. Analogous to buying an IKEA furniture, given a set of 3D parts that can assemble a single shape, an intelligent agent needs to perceive the 3D part geometry, reason to propose pose estimations for the inpu…

Cited by 100SourcePDFScholar
2019

RainFlow: Optical Flow Under Rain Streaks and Rain Veiling Effect

ICCV 2019poster

Optical flow in heavy rainy scenes is challenging due to the presence of both rain steaks and rain veiling effect, which break the existing optical flow constraints. Concerning this, we propose a deep-learning based optical flow method designed to handle heavy rain. We introduce a feature multiplier…

Cited by 43PDFScholar
2018

Decouple Learning for Parameterized Image Operators

ECCV 2018poster

Many different deep networks have been used to approximate, accelerate or improve traditional image operators, such as image smoothing, super-resolution and denoising. Among these traditional operators, many contain parameters which need to be tweaked to obtain the satisfactory results, which we ref…

2017

A Generic Deep Architecture for Single Image Reflection Removal and Image Smoothing

ICCV 2017poster

This paper proposes a deep neural network structure that exploits edge information in addressing representative low-level vision tasks such as layer separation and image filtering. Unlike most other deep learning strategies applied in this context, our approach tackles these challenging problems by…

Cited by 379PDFScholar