← Search

Mingming He

7 accepted papers

2026

Lighting in Motion: Spatiotemporal HDR Lighting Estimation

CVPR 2026

We present Lighting in Motion (LiMo), a diffusion-based approach to spatiotemporal lighting estimation. LiMo targets both realistic high-frequency detail prediction and accurate illuminance estimation. To account for both, we propose generating a set of mirrored and diffuse spheres at different expo

Cited by 0SourceScholar
2025

Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

CVPR 2025poster

Generative modeling aims to transform random noise into structured outputs. In this work, we enhance video diffusion models by allowing motion control via structured latent noise sampling. This is achieved by just a change in data: we pre-process training videos to yield structured noise. Consequent…

2025

Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset

CVPR 2025poster

Video portrait relighting remains challenging because the results need to be both photorealistic and temporally stable.This typically requires a strong model design that can capture complex facial reflections as well as intensive training on a high-quality paired video dataset, such as dynamic one-l…

Cited by 1SourcePDFScholar
2023

AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control

ICCV 2023poster

Neural implicit fields are powerful for representing 3D scenes and generating high-quality novel views, but it remains challenging to use such implicit representations for creating a 3D human avatar with a specific identity and artistic style that can be easily animated. Our proposed method, AvatarC…

Cited by 81PDFcodeScholar
2022

CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields

CVPR 2022poster

We present CLIP-NeRF, a multi-modal 3D object manipulation method for neural radiance fields (NeRF). By leveraging the joint language-image embedding space of the recent Contrastive Language-Image Pre-Training (CLIP) model, we propose a unified framework that allows manipulating NeRF in a user-frien…

Cited by 458PDFcodeScholar
2021

DisUnknown: Distilling Unknown Factors for Disentanglement Learning

ICCV 2021poster

Disentangling data into interpretable and independent factors is critical for controllable generation tasks. With the availability of labeled data, supervision can help enforce the separation of specific factors as expected. However, it is often expensive or even impossible to label every single fac…

Cited by 7PDFcodeScholar
2019

Deep Exemplar-Based Video Colorization

CVPR 2019poster

This paper presents the first end-to-end network for exemplar-based video colorization. The main challenge is to achieve temporal consistency while remaining faithful to the reference style. To address this issue, we introduce a recurrent framework that unifies the semantic correspondence and color…

Cited by 261PDFcodeScholar