← Search

Timo Bolkart

26 accepted papers

2026

Feed-forward Gaussian Registration for Head Avatar Creation and Editing

CVPR 2026

We present MATCH (Multi-view Avatars from Topologically Corresponding Heads), a multi-view Gaussian registration method for high-quality head avatar creation and editing. State-of-the-art multi-view head avatars require time-consuming head tracking, which is followed by an expensive avatar optimizat

Cited by 0SourcecodeScholar
2026

Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence

CVPR 2026

Recent frameworks like ToFu and TEMPEH provide an automated alternative to classical registration pipelines by predicting 3D meshes in dense semantic correspondence directly from calibrated multi-view images. However, these learning-based methods rely on the slow, manual registration pipelines they

Cited by 0SourcecodeScholar
2025

OFER: Occluded Face Expression Reconstruction

CVPR 2025poster

Reconstructing 3D face models from a single image is an inherently ill-posed problem, which becomes even more challenging in the presence of occlusions. In addition to fewer available observations, occlusions introduce an extra source of ambiguity where multiple reconstructions can be equally valid.…

Cited by 0SourcePDFScholar
2025

Synthetic Prior for Few-Shot Drivable Head Avatar Inversion

CVPR 2025poster

We present SynShot, a novel method for the few-shot inversion of a drivable head avatar based on a synthetic prior. We tackle three major challenges. First, training a controllable 3D generative network requires a large number of diverse sequences, for which pairs of images and high-quality tracked…

Cited by 1SourcePDFScholar
2024

3D Facial Expressions through Analysis-by-Neural-Synthesis

CVPR 2024poster

While existing methods for 3D face reconstruction from in-the-wild images excel at recovering the overall face shape they commonly miss subtle extreme asymmetric or rarely observed expressions. We improve upon these methods with SMIRK (Spatial Modeling for Image-based Reconstruction of Kinesics) whi…

2024

Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

CVPR 2024poster

Existing methods for synthesizing 3D human gestures from speech have shown promising results but they do not explicitly model the impact of emotions on the generated gestures. Instead these methods directly output animations from speech without control over the expressed emotion. To address this lim…

2024

SCULPT: Shape-Conditioned Unpaired Learning of Pose-dependent Clothed and Textured Human Meshes

CVPR 2024poster

We present SCULPT a novel 3D generative model for clothed and textured 3D meshes of humans. Specifically we devise a deep neural network that learns to represent the geometry and appearance distribution of clothed human bodies. Training such a model is challenging as datasets of textured 3D meshes f…

Cited by 5SourcePDFScholar
2024

Skin Tone Disentanglement in 2D Makeup Transfer With Graph Neural Networks

ICASSP 2024accepted

Makeup transfer involves transferring makeup from a reference image to a target image while maintaining the target’s identity. Existing methods, which use Generative Adversarial Networks, often transfer not just makeup but also the reference image’s skin tone. This limits their use to similar skin t…

Cited by 0SourceScholar
2023

Instant Multi-View Head Capture Through Learnable Registration

CVPR 2023poster

Existing methods for capturing datasets of 3D heads in dense semantic correspondence are slow and commonly address the problem in two separate steps; multi-view stereo (MVS) reconstruction followed by non-rigid registration. To simplify this process, we introduce TEMPEH (Towards Estimation of 3D Mes…

2022

SUPR: A Sparse Unified Part-Based Human Representation

ECCV 2022poster

"Statistical 3D shape models of the head, hands, and full body are widely used in computer vision and graphics. Despite their wide use, we show that existing models of the head and hands fail to capture the full range of motion for these parts. Moreover, existing work largely ignores the feet, which…

2022

Towards Racially Unbiased Skin Tone Estimation via Scene Disambiguation

ECCV 2022poster

"Virtual facial avatars will play an increasingly important role in immersive communication, games and the metaverse, and it is therefore critical that they be inclusive. This requires accurate recovery of the albedo, regardless of age, sex, or ethnicity. While significant progress has been made on…

Cited by 35SourcePDFScholar
2021

Learning Realistic Human Reposing Using Cyclic Self-Supervision With 3D Shape, Pose, and Appearance Consistency

ICCV 2021poster

Synthesizing images of a person in novel poses from a single image is a highly ambiguous task. Most existing approaches require paired training images; i.e. images of the same person with the same clothing in different poses. However, obtaining sufficiently large datasets with paired data is challen…

Cited by 20PDFScholar
2021

Topologically Consistent Multi-View Face Inference Using Volumetric Sampling

ICCV 2021poster

High-fidelity face digitization solutions often combine multi-view stereo (MVS) techniques for 3D reconstruction and a non-rigid registration step to establish dense correspondence across identities and expressions. A common problem is the need for manual clean-up after the MVS step, as 3D scans are…

Cited by 26PDFcodeScholar
2020

Monocular Expressive Body Regression through Body-Driven Attention

ECCV 2020poster

To understand how people look, interact, or perform tasks, we need to quickly and accurately capture their 3D body, face, and hands together from an RGB image. Most existing methods focus only on parts of the body. A few recent approaches reconstruct full expressive 3D humans from images using 3D bo…

2020

STAR: Sparse Trained Articulated Human Body Regressor

ECCV 2020poster

The SMPL body model is widely used for the estimation, synthesis, and analysis of 3D human pose and shape. While popular, we show that SMPL has several limitations and introduce STAR, which is quantitatively and qualitatively superior to SMPL. First, SMPL has a huge number of parameters resulting fr…

2019

Capture, Learning, and Synthesis of 3D Speaking Styles

CVPR 2019poster

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we introduce a unique 4D face dataset with about 29 minutes of 4D…

Cited by 426PDFcodeScholar
2019

Expressive Body Capture: 3D Hands, Face, and Body From a Single Image

CVPR 2019oral

To facilitate the analysis of human actions, interactions and emotions, we compute a 3D model of human body pose, hand pose, and facial expression from a single monocular image. To achieve this, we use thousands of 3D scans to train a new, unified, 3D model of the human body, SMPL-X, that extends SM…

Cited by 2080PDFcodeScholar
2019

Learning to Regress 3D Face Shape and Expression From an Image Without 3D Supervision

CVPR 2019poster

The estimation of 3D face shape from a single image must be robust to variations in lighting, head pose, expression, facial hair, makeup, and occlusions. Robustness requires a large training set of in-the-wild images, which by construction, lack ground truth 3D shape. To train a network without any…

Cited by 363PDFScholar
2018

Generating 3D Faces using Convolutional Mesh Autoencoders

ECCV 2018poster

Learned 3D representations of human faces are useful for computer vision problems such as 3D face tracking and reconstruction from images, as well as graphics applications such as character generation and animation. Traditional models learn a latent representation of a face using linear subspaces or…