← Search

Xuaner Zhang

14 accepted papers

2025

Instruction-based Image Manipulation by Watching How Things Move

CVPR 2025highlight

This paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation models. Video frames inherently preserve the identity of subjects…

Cited by 3SourcePDFScholar
2025

LEDiff: Latent Exposure Diffusion for HDR Generation

CVPR 2025poster

While consumer displays increasingly support more than 10 stops of dynamic range, most image assets -- such as internet photographs and generative AI content -- remain limited to 8-bit low dynamic range (LDR), constraining their utility across high dynamic range (HDR) applications. Currently, no gen…

Cited by 0SourcePDFScholar
2024

COMPOSE: Comprehensive Portrait Shadow Editing

ECCV 2024poster

"Existing portrait relighting methods struggle with precise control over facial shadows, particularly when faced with challenges such as handling hard shadows from directional light sources or adjusting shadows while remaining in harmony with existing lighting conditions. In many situations, complet…

Cited by 3SourcePDFScholar
2024

Dr. Bokeh: DiffeRentiable Occlusion-aware Bokeh Rendering

CVPR 2024poster

Bokeh is widely used in photography to draw attention to the subject while effectively isolating distractions in the background. Computational methods can simulate bokeh effects without relying on a physical camera lens but the inaccurate lens modeling in existing filtering-based methods leads to ar…

Cited by 8SourcePDFScholar
2024

Explorative Inbetweening of Time and Space

ECCV 2024poster

"We introduce bounded generation as a generalized task to control video generation to synthesize arbitrary camera and subject motion based only on a given start and end frame. Our objective is to fully leverage the inherent generalization capability of an image-to-video model without additional trai…

Cited by 12SourcePDFScholar
2024

Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image

CVPR 2024poster

At the core of portrait photography is the search for ideal lighting and viewpoint. The process often requires advanced knowledge in photography and an elaborate studio setup. In this work we propose Holo-Relighting a volumetric relighting method that is capable of synthesizing novel viewpoints and…

Cited by 12SourcePDFScholar
2023

Automatic High Resolution Wire Segmentation and Removal

CVPR 2023poster

Wires and powerlines are common visual distractions that often undermine the aesthetics of photographs. The manual process of precisely segmenting and removing them is extremely tedious and may take up to hours, especially on high-resolution photos where wires may span the entire space. In this pape…

2023

DiffusionRig: Learning Personalized Priors for Facial Appearance Editing

CVPR 2023poster

We address the problem of learning person-specific facial priors from a small number (e.g., 20) of portrait photos of the same person. This enables us to edit this specific person's facial appearance, such as expression and lighting, while preserving their identity and high-frequency facial details.…

2023

LightPainter: Interactive Portrait Relighting With Freehand Scribble

CVPR 2023poster

Recent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-b…

Cited by 14SourcePDFScholar
2023

SunStage: Portrait Reconstruction and Relighting Using the Sun as a Light Stage

CVPR 2023poster

A light stage uses a series of calibrated cameras and lights to capture a subject's facial appearance under varying illumination and viewpoint. This captured information is crucial for facial reconstruction and relighting. Unfortunately, light stages are often inaccessible: they are expensive and re…

Cited by 29SourcePDFScholar
2022

The Implicit Values of a Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement

CVPR 2022oral

Modern smartphones can continuously stream multi-megapixel RGB images at 60Hz, synchronized with high-quality 3D pose information and low-resolution LiDAR-driven depth estimates. During a snapshot photograph, the natural unsteadiness of the photographer's hands offers millimeter-scale variation in c…

Cited by 17PDFcodeScholar