← Search

Zhihao Xia

17 accepted papers

2026

RhoMorph: Rhombus-Shaped Modular Robots for Stable, Medium-Independent Reconfiguration Motion

ICRA 2026poster

In this paper, we present RhoMorph, a novel deformable planar lattice modular self-reconfigurable robot (MSRR) with a rhombus shaped module. Each module consists of a parallelogram skeleton with a single centrally mounted actuator that enables folding and unfolding along its diagonal. The core desig…

Cited by 0Scholar
2026

Self-Reconfiguration Planning for Deformable Quadrilateral Modular Robots

RA-L 2026

While deformable modular self reconfigurable robots offer enhanced reconfiguration flexibility, strict kinematic constraints present complex self reconfiguration planning challenges. This letter presents a novel self-reconfiguration planning algorithm for deformable quadrilateral MSRRs. The method f

Cited by 0SourceScholar
2025

Classic Video Denoising in a Machine Learning World: Robust, Fast, and Controllable

CVPR 2025poster

Denoising is a crucial step in many video processing pipelines such as in interactive editing, where high quality, speed, and user control are essential. While recent approaches achieve significant improvements in denoising quality by leveraging deep learning, they are prone to unexpected failures d…

Cited by 0SourcePDFScholar
2025

Instruction-based Image Manipulation by Watching How Things Move

CVPR 2025highlight

This paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation models. Video frames inherently preserve the identity of subjects…

Cited by 3SourcePDFScholar
2025

LEDiff: Latent Exposure Diffusion for HDR Generation

CVPR 2025poster

While consumer displays increasingly support more than 10 stops of dynamic range, most image assets -- such as internet photographs and generative AI content -- remain limited to 8-bit low dynamic range (LDR), constraining their utility across high dynamic range (HDR) applications. Currently, no gen…

Cited by 0SourcePDFScholar
2024

Explorative Inbetweening of Time and Space

ECCV 2024poster

"We introduce bounded generation as a generalized task to control video generation to synthesize arbitrary camera and subject motion based only on a given start and end frame. Our objective is to fully leverage the inherent generalization capability of an image-to-video model without additional trai…

Cited by 12SourcePDFScholar
2023

DiffusionRig: Learning Personalized Priors for Facial Appearance Editing

CVPR 2023poster

We address the problem of learning person-specific facial priors from a small number (e.g., 20) of portrait photos of the same person. This enables us to edit this specific person's facial appearance, such as expression and lighting, while preserving their identity and high-frequency facial details.…

2023

Semi-Supervised Parametric Real-World Image Harmonization

CVPR 2023poster

Learning-based image harmonization techniques are usually trained to undo synthetic global transformations, applied to a masked foreground in a single ground truth photo. This simulated data does not model many important appearance mismatches (illumination, object boundaries, etc.) between foregroun…

2022

The Implicit Values of a Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement

CVPR 2022oral

Modern smartphones can continuously stream multi-megapixel RGB images at 60Hz, synchronized with high-quality 3D pose information and low-resolution LiDAR-driven depth estimates. During a snapshot photograph, the natural unsteadiness of the photographer's hands offers millimeter-scale variation in c…

Cited by 17PDFcodeScholar
2021

Deep Denoising of Flash and No-Flash Pairs for Photography in Low-Light Environments

CVPR 2021poster

We introduce a neural network-based method to denoise pairs of images taken in quick succession in low-light environments, with and without a flash. Our goal is to produce a high-quality rendering of the scene that preserves the color and mood from the ambient illumination of the noisy no-flash imag…

Cited by 25PDFScholar
2020

Basis Prediction Networks for Effective Burst Denoising With Large Kernels

CVPR 2020poster

Bursts of images exhibit significant self-similarity across both time and space. This motivates a representation of the kernels as linear combinations of a small set of basis elements. To this end, we introduce a novel basis prediction network that, given an input burst, predicts a set of global bas…

Cited by 86PDFScholar