← Search

Chen Cao

12 accepted papers

2026

GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

CVPR 2026

Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error accumulation, or they do not explicitly leverage 3D geometry du

Cited by 0SourceScholar
2026

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

CVPR 2026

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and

Cited by 0SourcecodeScholar
2025

Advancing Audio-Based Text Generation with Imbalance Preference Optimization

AAAI 2025technical

Human feedback in generative systems is a highly active frontier of research that aims to improve the quality of generated content and align it with subjective preferences. Existing efforts predominantly focus on text-only large language models (LLMs) or text-based image generation, while cross-moda…

Cited by 0SourcePDFScholar
2025

LUCAS: Layered Universal Codec Avatars

CVPR 2025poster

Photorealistic 3D head avatar reconstruction faces critical challenges in modeling dynamic face-hair interactions and achieving cross-identity generalization, particularly during expressions and head movements. We present LUCAS, a novel Universal Prior Model (UPM) for codec avatar modeling that dise…

Cited by 0SourcePDFScholar
2025

Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior

CVPR 2025poster

We present Vid2Avatar-Pro, a method to create photorealistic and animatable 3D human avatars from monocular in-the-wild videos. Building a high-quality avatar that supports animation with diverse poses from a monocular video is challenging because the observation of pose diversity and view points is…

Cited by 0SourcePDFScholar
2024

Bridging the Gap: Studio-like Avatar Creation from a Monocular Phone Capture

ECCV 2024oral

"Creating photorealistic avatars for individuals traditionally involves extensive capture sessions with complex and expensive studio devices like the LightStage system. While recent strides in neural representations have enabled the generation of photorealistic and animatable 3D avatars from quick p…

Cited by 0SourcePDFScholar
2024

URHand: Universal Relightable Hands

CVPR 2024poster

Existing photorealistic relightable hand models require extensive identity-specific observations in different views poses and illuminations and face challenges in generalizing to natural illuminations and novel identities. To bridge this gap we present URHand the first universal relightable hand mod…

Cited by 11SourcePDFScholar
2023

NeuWigs: A Neural Dynamic Model for Volumetric Hair Capture and Animation

CVPR 2023poster

The capture and animation of human hair are two of the major challenges in the creation of realistic avatars for the virtual reality. Both problems are highly challenging, because hair has complex geometry and appearance, as well as exhibits challenging motion. In this paper, we present a two-stage…

Cited by 16SourcePDFScholar
2021

High-Fidelity Face Tracking for AR/VR via Deep Lighting Adaptation

CVPR 2021poster

3D video avatars can empower virtual communications by providing compression, privacy, entertainment, and a sense of presence in AR/VR. Best 3D photo-realistic AR/VR avatars driven by video, that can minimize uncanny effects, rely on person-specific models. However, existing person-specific photo-re…

Cited by 29PDFScholar
2020

Fully Convolutional Mesh Autoencoder using Efficient Spatially Varying Kernels

NeurIPS 2020poster

Learning latent representations of registered meshes is useful for many 3D tasks. Techniques have recently shifted to neural mesh autoencoders. Although they demonstrate higher precision than traditional methods, they remain unable to capture fine-grained deformations. Furthermore, these methods can…

Cited by 97SourcePDFScholar