← Search

Yu-Lun Liu

37 accepted papers

2026

Pantheon360: Taming Digital Twin Generation via 3D-Aware 360deg Video Diffusion

CVPR 2026

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency--constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view tra

Cited by 0SourceScholar
2026

PhaSR: Generalized Image Shadow Removal with Physically Aligned Priors

CVPR 2026

Shadow removal under diverse lighting conditions requires disentangling illumination from intrinsic reflectance--a challenge compounded when physical priors are not properly aligned. We propose PhaSR (Physically Aligned Shadow Removal), addressing this through dual-level prior alignment to enable ro

Cited by 0SourcecodeScholar
2026

Reflection Separation from a Single Image via Joint Latent Diffusion

CVPR 2026

Single-image reflection separation is highly challenging under extreme conditions like glare or weak reflections. Existing methods often struggle to recover both layers in glare or weak-reflection scenarios because of insufficient information. This paper presents a diffusion model explicitly fine-tu

Cited by 0SourcecodeScholar
2026

ReflexSplit: Single Image Reflection Separation via Layer Fusion-Separation

CVPR 2026

Single Image Reflection Separation (SIRS) disentangles mixed images into transmission and reflection layers. Existing methods suffer from transmission-reflection confusion under nonlinear mixing, particularly in deep decoder layers, due to implicit fusion mechanisms and inadequate multi-scale coordi

Cited by 0SourcecodeScholar
2025

3D Gaussian Splatting with Grouped Uncertainty for Unconstrained Images

ICASSP 2025accepted

3D Gaussian Splatting (3DGS) [1] is a promising method for 3D reconstruction and novel view synthesis. However, training it with unconstrained images presents challenges due to transient objects that cause undesired floaters and ghosting artifacts. Although related works using Neural Radiance Fields…

Cited by 0SourceScholar
2025

AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene Inpainting

CVPR 2025poster

Three-dimensional scene inpainting is crucial for applications from virtual reality to architectural visualization, yet existing methods struggle with view consistency and geometric accuracy in 360deg unbounded scenes. We present AuraFusion360, a novel reference-based method that enables high-qualit…

Cited by 0SourcePDFScholar
2025

CAT-3DGS: A Context-Adaptive Triplane Approach to Rate-Distortion-Optimized 3DGS Compression

ICLR 2025poster

3D Gaussian Splatting (3DGS) has recently emerged as a promising 3D representation. Much research has been focused on reducing its storage requirements and memory footprint. However, the needs to compress and transmit the 3DGS representation to the remote side are overlooked. This new application ca…

Cited by 3SourcePDFScholar
2025

DeNVeR: Deformable Neural Vessel Representations for Unsupervised Video Vessel Segmentation

CVPR 2025poster

This paper presents Deformable Neural Vessel Representations (DeNVeR), an unsupervised approach for vessel segmentation in X-ray angiography videos without annotated ground truth. DeNVeR utilizes optical flow and layer separation techniques, enhancing segmentation accuracy and adaptability through t…

2025

FIPER: Factorized Features for Robust Image Super-Resolution and Compression

NeurIPS 2025poster

In this work, we propose using a unified representation, termed **Factorized Features**, for low-level vision tasks, where we test on **Single Image Super-Resolution (SISR)** and **Image Compression**. Motivated by the shared principles between these tasks, they require recovering and preserving fin…

Cited by 0SourceScholar
2025

FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priors

CVPR 2025poster

Neural Radiance Fields (NeRF) face significant challenges in extreme few-shot scenarios, primarily due to overfitting and long training times. Existing methods, such as FreeNeRF and SparseNeRF, use frequency regularization or pre-trained priors but struggle with complex scheduling and bias. We intro…

Cited by 0SourcePDFScholar
2025

GCC: Generative Color Constancy via Diffusing a Color Checker

CVPR 2025poster

Color constancy methods often struggle to generalize across different camera sensors due to varying spectral sensitivities. We present GCC, which leverages diffusion models to inpaint color checkers into images for illumination estimation. Our key innovations include (1) a single-step deterministic…

Cited by 0SourcePDFScholar
2025

LightsOut: Diffusion-based Outpainting for Enhanced Lens Flare Removal

ICCV 2025poster

Lens flare significantly degrades image quality, impacting critical computer vision tasks like object detection and autonomous driving. Recent Single Image Flare Removal (SIFR) methods perform poorly when off-frame light sources are incomplete or absent. We propose LightsOut, a diffusion-based outpa…

Cited by 0SourcePDFScholar
2025

LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

ICCV 2025poster

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift, inaccurate geometry initialization, and severe memory limitatio…

2025

OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations

ICCV 2025poster

Open-vocabulary (OV) 3D object detection is an emerging field, yet its exploration through image-based methods remains limited compared to 3D point cloud-based methods. We introduce OpenM3D, a novel open-vocabulary multi-view indoor 3D object detector trained without human annotations. In particular…

Cited by 0SourcePDFScholar
2025

See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation

CoRL 2025poster

We present See, Point, Fly (SPF), a training-free aerial vision-and-language navigation (AVLN) framework built atop vision-language models (VLMs). SPF is capable of navigating to any goal based on any type of free-form instructions in any kind of environment. In contrast to existing VLM-based approa…

Cited by 0SourceScholar
2025

SpectroMotion: Dynamic 3D Reconstruction of Specular Scenes

CVPR 2025poster

We present SpectroMotion, a novel approach that combines 3D Gaussian Splatting (3DGS) with physically-based rendering (PBR) and deformation fields to reconstruct dynamic specular scenes. Previous methods extending 3DGS to model dynamic scenes have struggled to represent specular surfaces accurately.…

2025

StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided Illusions

ICCV 2025poster

3D scene representation methods like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have significantly advanced novel view synthesis. As these methods become prevalent, addressing their vulnerabilities becomes critical. We analyze 3DGS robustness against image-level poisoning attacks…

Cited by 0SourcePDFScholar
2024

Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data Augmentation

NeurIPS 2024poster

Accurately estimating depth in 360-degree imagery is crucial for virtual reality, autonomous navigation, and immersive media applications. Existing depth estimation methods designed for perspective-view imagery fail when applied to 360-degree images due to different camera projections and distortion…

Cited by 4SourcePDFScholar
2024

Dual Associated Encoder for Face Restoration

ICLR 2024poster

Restoring facial details from low-quality (LQ) images has remained challenging due to the nature of the problem caused by various degradations in the wild. The codebook prior has been proposed to address the ill-posed problems by leveraging an autoencoder and learned codebook of high-quality (HQ) f…

2024

GenRC: Generative 3D Room Completion from Sparse Image Collections

ECCV 2024poster

"Sparse RGBD scene completion is a challenging task especially when considering consistent textures and geometries throughout the entire scene. Different from existing solutions that rely on human-designed text prompts or predefined camera trajectories, we propose , an automated training-free pipeli…

2024

HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse Poses

CVPR 2024poster

We present HumanNeRF-SE a simple yet effective method that synthesizes diverse novel pose images with simple input. Previous HumanNeRF works require a large number of optimizable parameters to fit the human images. Instead we reload these approaches by combining explicit and implicit human represent…

Cited by 4SourcePDFScholar
2024

Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation

CVPR 2024poster

This paper addresses text-supervised semantic segmentation aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs without dense annotations. Existing methods have demonstrated that contrastive learning on image-text pairs effectively alig…

2024

Improving Robustness for Joint Optimization of Camera Pose and Decomposed Low-Rank Tensorial Radiance Fields

AAAI 2024technical

In this paper, we propose an algorithm that allows joint refinement of camera pose and scene geometry represented by decomposed low-rank tensor, using only 2D images as supervision. First, we conduct a pilot study based on a 1D signal and relate our findings to 3D scenarios, where the naive joint…

2024

NaRCan: Natural Refined Canonical Image with Integration of Diffusion Prior for Video Editing

NeurIPS 2024poster

We propose a video editing framework, NaRCan, which integrates a hybrid deformation field and diffusion prior to generate high-quality natural canonical images to represent the input video. Our approach utilizes homography to model global motion and employs multi-layer perceptrons (MLPs) to capture…

2024

ReF-LDM: A Latent Diffusion Model for Reference-based Face Image Restoration

NeurIPS 2024poster

While recent works on blind face image restoration have successfully produced impressive high-quality (HQ) images with abundant details from low-quality (LQ) input images, the generated content may not accurately reflect the real appearance of a person. To address this problem, incorporating well-sh…

Cited by 0SourcePDFScholar
2023

ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection

ICCV 2023poster

We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-…

Cited by 11PDFcodeScholar
2023

Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction

ICCV 2023poster

Deep learning is commonly used to produce impressive results in reconstructing HDR images from LDR images. LDR stack-based methods are used for single-image HDR reconstruction, generating an HDR image from a deep learning generated LDR stack. However, current methods generate the LDR stack with pred…

Cited by 10PDFScholar
2023

Progressively Optimized Local Radiance Fields for Robust View Synthesis

CVPR 2023poster

We present an algorithm for reconstructing the radiance field of a large-scale scene from a single casually captured video. The task poses two core challenges. First, most existing radiance field reconstruction approaches rely on accurate pre-estimated camera poses from Structure-from-Motion algorit…

Cited by 106SourcePDFScholar
2023

Robust Dynamic Radiance Fields

CVPR 2023poster

Dynamic radiance field reconstruction methods aim to model the time-varying structure and appearance of a dynamic scene. Existing methods, however, assume that accurate camera poses can be reliably estimated by Structure from Motion (SfM) algorithms. These methods, thus, are unreliable as SfM algori…

2022

Denoising Likelihood Score Matching for Conditional Score-based Data Generation

ICLR 2022poster

Many existing conditional score-based data generation methods utilize Bayes' theorem to decompose the gradients of a log posterior density into a mixture of scores. These methods facilitate the training procedure of conditional score models, as a mixture of scores can be separately estimated using a…

2021

Bridging Unsupervised and Supervised Depth From Focus via All-in-Focus Supervision

ICCV 2021poster

Depth estimation is a long-lasting yet important task in computer vision. Most of the previous works try to estimate depth from input images and assume images are all-in-focus (AiF), which is less common in real-world applications. On the other hand, a few works take defocus blur into account and co…

Cited by 29PDFcodeScholar
2021

Hybrid Neural Fusion for Full-Frame Video Stabilization

ICCV 2021poster

Existing video stabilization methods often generate visible distortion or require aggressive cropping of frame boundaries, resulting in smaller field of views. In this work, we present a frame synthesis algorithm to achieve full-frame video stabilization. We first estimate dense warp fields from nei…

Cited by 59PDFcodeScholar
2020

Learning Camera-Aware Noise Models

ECCV 2020poster

Modeling imaging sensor noise is a fundamental problem for image processing and computer vision applications. While most previous works adopt statistical noise models, real-world noise is far more complicated and beyond what these models can describe. To tackle this issue, we propose a data-driven a…

2020

Single-Image HDR Reconstruction by Learning to Reverse the Camera Pipeline

CVPR 2020poster

Recovering a high dynamic range (HDR) image from a single low dynamic range (LDR) input image is challenging due to missing details in under-/over-exposed regions caused by quantization and saturation of camera sensors. In contrast to existing learning-based methods, our core idea is to incorporate…

Cited by 307PDFcodeScholar