← Search

Fangzhou Mu

11 accepted papers

2025

Recovering Parametric Scenes from Very Few Time-of-Flight Pixels

ICCV 2025poster

We aim to recover the geometry of 3D parametric scenes using very few depth measurements from low-cost, commercially available time-of-flight sensors. These sensors offer very low spatial resolution (i.e., a single pixel), but image a wide field-of-view per pixel and capture detailed time-of-flight…

Cited by 0SourcePDFScholar
2024

Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance

NeurIPS 2024poster

Recent controllable generation approaches such as FreeControl and Diffusion Self-Guidance bring fine-grained spatial and appearance control to text-to-image (T2I) diffusion models without training auxiliary modules. However, these methods optimize the latent embedding for each type of score function…

2024

FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition

CVPR 2024poster

Recent approaches such as ControlNet offer users fine-grained spatial control over text-to-image (T2I) diffusion models. However auxiliary modules have to be trained for each spatial condition type model architecture and checkpoint putting them at odds with the diverse intents and preferences a huma…

2024

Towards 3D Vision with Low-Cost Single-Photon Cameras

CVPR 2024poster

We present a method for reconstructing 3D shape of arbitrary Lambertian objects based on measurements by miniature energy-efficient low-cost single-photon cameras. These cameras operating as time resolved image sensors illuminate the scene with a very fast pulse of diffuse light and record the shape…

Cited by 10SourcePDFScholar
2024

Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning

ICLR 2024poster

Foundation models have emerged as a powerful tool for many AI problems. Despite the tremendous success of foundation models, effective adaptation to new tasks, particularly those with limited labels, remains an open question and lacks theoretical understanding. An emerging solution with recent su…

2023

GLIGEN: Open-Set Grounded Text-to-Image Generation

CVPR 2023poster

Large-scale text-to-image diffusion models have made amazing advances. However, the status quo is to use text input alone, which can impede controllability. In this work, we propose GLIGEN: Open-Set Grounded Text-to-Image Generation, a novel approach that builds upon and extends the functionality of…

2023

Learned Compressive Representations for Single-Photon 3D Imaging

ICCV 2023poster

Single-photon 3D cameras can record the time-of-arrival of billions of photons per second with picosecond accuracy. One common approach to summarize the photon data stream is to build a per-pixel timestamp histogram, resulting in a 3D histogram tensor that encodes distances along the time axis. As t…

Cited by 4PDFScholar
2022

3D Photo Stylization: Learning To Generate Stylized Novel Views From a Single Image

CVPR 2022oral

Visual content creation has spurred a soaring interest given its applications in mobile photography and AR / VR. Style transfer and single-image 3D photography as two representative tasks have so far evolved independently. In this paper, we make a connection between the two, and address the challeng…

Cited by 60PDFScholar
2022

SmartAdapt: Multi-Branch Object Detection Framework for Videos on Mobiles

CVPR 2022poster

Several recent works seek to create lightweight deep networks for video object detection on mobiles. We observe that many existing detectors, previously deemed computationally costly for mobiles, intrinsically support adaptive inference, and offer a multi-branch object detection framework (MBODF). H…

Cited by 15PDFScholar