← Search

Thiemo Alldieck

13 accepted papers

2025

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

CVPR 2025poster

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-to-3d-motion diffusion model, and 2) a novel diffusion-based architecture that a…

Cited by 27SourcePDFScholar
2024

DiffHuman: Probabilistic Photorealistic 3D Reconstruction of Humans

CVPR 2024poster

We present DiffHuman a probabilistic method for photorealistic 3D human reconstruction from a single RGB image. Despite the ill-posed nature of this problem most methods are deterministic and output a single solution often resulting in a lack of geometric detail and blurriness in unseen or uncertain…

Cited by 5SourcePDFScholar
2024

Instant 3D Human Avatar Generation using Image Diffusion Models

ECCV 2024poster

"We present , a method for fast, high quality 3D human avatar generation from different input modalities, such as images and text prompts and with control over the generated pose and shape. The common theme is the use of diffusion-based image generation networks that are specialized for each particu…

Cited by 5SourcePDFScholar
2023

DreamHuman: Animatable 3D Avatars from Text

NeurIPS 2023spotlight

We present \emph{DreamHuman}, a method to generate realistic animatable 3D human avatar models entirely from textual descriptions. Recent text-to-3D methods have made considerable strides in generation, but are still lacking in important aspects. Control and often spatial resolution remain limited,…

Cited by 96SourcePDFScholar
2023

Structured 3D Features for Reconstructing Controllable Avatars

CVPR 2023poster

We introduce Structured 3D Features, a model based on a novel implicit 3D representation that pools pixel-aligned image features onto dense 3D points sampled from a parametric, statistical human mesh surface. The 3D points have associated semantics and can move freely in 3D space. This allows for op…

Cited by 39SourcePDFScholar
2022

Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing

CVPR 2022poster

We present PHORHUM, a novel, end-to-end trainable, deep neural network methodology for photorealistic 3D human reconstruction given just a monocular RGB image. Our pixel-aligned method estimates detailed 3D geometry and, for the first time, the unshaded surface color together with the scene illumina…

Cited by 161PDFScholar
2021

H-NeRF: Neural Radiance Fields for Rendering and Temporal Reconstruction of Humans in Motion

NeurIPS 2021spotlight

We present neural radiance fields for rendering and temporal (4D) reconstruction of humans in motion (H-NeRF), as captured by a sparse set of cameras or even from a monocular video. Our approach combines ideas from neural scene representation, novel-view synthesis, and implicit statistical geometric…

Cited by 205SourcePDFScholar
2021

imGHUM: Implicit Generative Models of 3D Human Shape and Articulated Pose

ICCV 2021poster

We present imGHUM, the first holistic generative model of 3D human shape and articulated pose, represented as a signed distance function. In contrast to prior work, we model the full human body implicitly as a function zero-level-set and without the use of an explicit template mesh. We propose a nov…

Cited by 121PDFcodeScholar
2020

Implicit Functions in Feature Space for 3D Shape Reconstruction and Completion

CVPR 2020poster

While many works focus on 3D reconstruction from images, in this paper, we focus on 3D shape reconstruction and completion from a variety of 3D inputs, which are deficient in some respect: low and high resolution voxels, sparse and dense point clouds, complete or incomplete. Processing of such 3D in…

Cited by 578PDFcodeScholar
2019

Learning to Reconstruct People in Clothing From a Single RGB Camera

CVPR 2019poster

We present Octopus, a learning-based model to infer the personalized 3D shape of people from a few frames (1-8) of a monocular video in which the person is moving with a reconstruction accuracy of 4 to 5mm, while being orders of magnitude faster than previous methods. From semantic segmentation imag…

Cited by 380PDFcodeScholar
2019

Tex2Shape: Detailed Full Human Body Geometry From a Single Image

ICCV 2019poster

We present a simple yet effective method to infer detailed full human body shape from only a single photograph. Our model can infer full-body shape including face, hair, and clothing including wrinkles at interactive frame-rates. Results feature details even on parts that are occluded in the input i…

Cited by 372PDFcodeScholar
2018

Video Based Reconstruction of 3D People Models

CVPR 2018poster

This paper describes how to obtain accurate 3D body models and texture of arbitrary people from a single, monocular video in which a person is moving. Based on a parametric body model, we present a robust processing pipeline achieving 3D model fits with 5mm accuracy also for clothed people. Our main…