← Search

Andrei Zanfir

14 accepted papers

2025

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

CVPR 2025poster

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-to-3d-motion diffusion model, and 2) a novel diffusion-based architecture that a…

Cited by 27SourcePDFScholar
2024

DiffHuman: Probabilistic Photorealistic 3D Reconstruction of Humans

CVPR 2024poster

We present DiffHuman a probabilistic method for photorealistic 3D human reconstruction from a single RGB image. Despite the ill-posed nature of this problem most methods are deterministic and output a single solution often resulting in a lack of geometric detail and blurriness in unseen or uncertain…

Cited by 5SourcePDFScholar
2023

DreamHuman: Animatable 3D Avatars from Text

NeurIPS 2023spotlight

We present \emph{DreamHuman}, a method to generate realistic animatable 3D human avatar models entirely from textual descriptions. Recent text-to-3D methods have made considerable strides in generation, but are still lacking in important aspects. Control and often spatial resolution remain limited,…

Cited by 96SourcePDFScholar
2023

Structured 3D Features for Reconstructing Controllable Avatars

CVPR 2023poster

We introduce Structured 3D Features, a model based on a novel implicit 3D representation that pools pixel-aligned image features onto dense 3D points sampled from a parametric, statistical human mesh surface. The 3D points have associated semantics and can move freely in 3D space. This allows for op…

Cited by 39SourcePDFScholar
2022

HUM3DIL: Semi-supervised Multi-modal 3D HumanPose Estimation for Autonomous Driving

CoRL 2022poster

Autonomous driving is an exciting new industry, posing important research questions. Within the perception module, 3D human pose estimation is an emerging technology, which can enable the autonomous vehicle to perceive and understand the subtle and complex behaviors of pedestrians. While hardware sy…

Cited by 32SourceScholar
2021

Neural Descent for Visual 3D Human Pose and Shape

CVPR 2021poster

We present deep neural network methodology to reconstruct the 3d pose and shape of people, including hand gestures and facial expression, given an input RGB image. We rely on a recently introduced, expressive full body statistical 3d human model, GHUM, trained end-to-end, and learn to reconstruct it…

Cited by 75PDFScholar
2021

THUNDR: Transformer-Based 3D Human Reconstruction With Markers

ICCV 2021poster

We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine the predictive power of model-free-output architectures and t…

Cited by 83PDFScholar
2020

GHUM & GHUML: Generative 3D Human Shape and Articulated Pose Models

CVPR 2020oral

We present a statistical, articulated 3D human shape modeling pipeline, within a fully trainable, modular, deep learning framework. Given high-resolution complete 3D body scans of humans, captured in various poses, together with additional closeups of their head and facial expressions, as well as ha…

Cited by 423PDFcodeScholar
2020

Weakly Supervised 3D Human Pose and Shape Reconstruction with Normalizing Flows

ECCV 2020poster

Monocular 3D human pose and shape estimation is challenging due to the many degrees of freedom of the human body and the difficulty to acquire training data for large-scale supervised learning in complex visual scenes where humans with diverse shape and appearance, appear against complex backgrounds…

Cited by 160SourcePDFScholar
2018

Deep Network for the Integrated 3D Sensing of Multiple People in Natural Images

NeurIPS 2018spotlight

We present MubyNet -- a feed-forward, multitask, bottom up system for the integrated localization, as well as 3d pose and shape estimation, of multiple people in monocular images. The challenge is the formal modeling of the problem that intrinsically requires discrete and continuous computation, e.g…

Cited by 170SourcePDFScholar
2018

Monocular 3D Pose and Shape Estimation of Multiple People in Natural Scenes - The Importance of Multiple Scene Constraints

CVPR 2018poster

Human sensing has greatly benefited from recent advances in deep learning, parametric human modeling, and large scale 2d and 3d datasets. However, existing 3d models make strong assumptions about the scene, considering either a single person per image, full views of the person, a simple background o…

Cited by 346SourcePDFScholar