← Search

Liao Shen

10 accepted papers

2026

BokehCrafter: Taming Video Diffusion Models for Controllable Bokeh Rendering

AAAI 2026technical

Bokeh is used in photography to emphasize the selected subject by smoothly blurring the out-of-focus region with appealing highlights. While recent advances have achieved impressive results in rendering realistic blur, existing frameworks typically rely on disparity maps and bokeh-relevant inputs (e

Cited by 0SourcePDFScholar
2026

BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow Matching

AAAI 2026technical

Bokeh rendering simulates the shallow depth-of-field effect in photography, enhancing visual aesthetics and guiding viewer attention to regions of interest. Although recent approaches perform well, rendering controllable bokeh without additional depth inputs remains a significant challenge. Existing

Cited by 0SourcePDFScholar
2026

Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization

CVPR 2026

Recent advances in image-to-video (I2V) generation have achieved remarkable progress in synthesizing high-quality, temporally coherent videos from static images. Among all the applications of I2V, human-centric video generation includes a large portion. However, existing I2V models encounter difficu

Cited by 0SourcecodeScholar
2025

CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching

CVPR 2025highlight

Depth estimation is a fundamental task in 3D vision. An ideal depth estimation model is expected to embrace meticulous detail, temporal consistency, and high efficiency. Although existing foundation models can perform well in certain specific aspects, most of them fall short of fulfilling all the ab…

Cited by 0SourcePDFScholar
2025

DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian Splatting

CVPR 2025poster

Recent advances in 3D Gaussian Splatting (3D-GS) have shown remarkable success in representing 3D scenes and generating high-quality, novel views in real-time. However, 3D-GS and its variants assume that input images are captured based on pinhole imaging and are fully in focus. This assumption limit…

2025

Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency

ICCV 2025poster

We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video datasets for expensive training, with limited generalization abili…

2025

MuGS: Multi-Baseline Generalizable Gaussian Splatting Reconstruction

ICCV 2025poster

We present Multi-Baseline Gaussian Splatting (MuGS), a generalized feed-forward approach for novel view synthesis that effectively handles diverse baseline settings, including sparse input views with both small and large baselines. Specifically, we integrate features from Multi-View Stereo (MVS) and…

2024

DreamMover: Leveraging the Prior of Diffusion Models for Image Interpolation with Large Motion

ECCV 2024poster

"We study the problem of generating intermediate images from image pairs with large motion while maintaining semantic consistency. Due to the large motion, the intermediate semantic information may be absent in input images. Existing methods either limit to small motion or focus on topologically sim…

2024

DyBluRF: Dynamic Neural Radiance Fields from Blurry Monocular Video

CVPR 2024poster

Recent advancements in dynamic neural radiance field methods have yielded remarkable outcomes. However these approaches rely on the assumption of sharp input images. When faced with motion blur existing dynamic NeRF methods often struggle to generate high-quality novel views. In this paper we propos…

Cited by 10SourcePDFScholar
2024

MVSGaussian: Fast Generalizable Gaussian Splatting Reconstruction from Multi-View Stereo

ECCV 2024poster

"We present MVSGaussian, a new generalizable 3D Gaussian representation approach derived from Multi-View Stereo (MVS) that can efficiently reconstruct unseen scenes. Specifically, 1) we leverage MVS to encode geometry-aware Gaussian representations and decode them into Gaussian parameters. 2) To fur…