← Search

Justus Thies

44 accepted papers

2026

CLUTCH: Contextualized Language model for Unlocking Text-Conditioned Hand motion modelling in the wild

ICLR 2026poster

Hands play a central role in daily life, yet modeling natural hand motions remains underexplored. Existing methods that tackle text-to-hand-motion generation or hand animation captioning rely on studio-captured datasets with limited actions and contexts, making them costly to scale to “in-the-wild”…

Cited by 0SourceScholar
2026

PhysHead: Simulation-Ready Gaussian Head Avatars

CVPR 2026

Realistic digital avatars require expressive and dynamic hair motion; however, most existing head avatar methods assume rigid hair movement. These methods often fail to disentangle hair from the head, representing it as a simple outer shell and failing to capture its natural volumetric behavior. In

Cited by 0SourcecodeScholar
2025

GaussianSpeech: Audio-Driven Personalized 3D Gaussian Avatars

ICCV 2025poster

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photorealistic and personalized multi-view consistent 3D human head avatars from spoken audio at real-time rendering rates. To capture the expressive and detailed nature of human heads, including skin…

2025

HairFree: Compositional 2D Head Prior for Text-Driven 360° Bald Texture Synthesis

NeurIPS 2025poster

Synthesizing high-quality 3D head textures is crucial for gaming, virtual reality, and digital humans. Achieving seamless 360° textures typically requires expensive multi-view datasets with precise tracking. However, traditional methods struggle without back-view data or precise geometry, especially…

Cited by 0SourceScholar
2025

Im2Haircut: Single-view Strand-based Hair Reconstruction for Human Avatars

ICCV 2025poster

We present a novel approach for 3D hair reconstruction from single photographs based on a global hair prior combined with local optimization. Capturing strand-based hair geometry from single photographs is challenging due to the variety and geometric complexity of hairstyles and the lack of ground t…

Cited by 0SourcePDFScholar
2025

Synthetic Prior for Few-Shot Drivable Head Avatar Inversion

CVPR 2025poster

We present SynShot, a novel method for the few-shot inversion of a drivable head avatar based on a synthetic prior. We tackle three major challenges. First, training a controllable 3D generative network requires a large number of diverse sequences, for which pairs of images and high-quality tracked…

Cited by 1SourcePDFScholar
2024

DPHMs: Diffusion Parametric Head Models for Depth-based Tracking

CVPR 2024poster

We introduce Diffusion Parametric Head Models (DPHMs) a generative model that enables robust volumetric head reconstruction and tracking from monocular depth sequences. While recent volumetric head models such as NPHMs can now excel in representing high-fidelity head geometries tracking and reconstr…

Cited by 6SourcePDFScholar
2024

DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis

CVPR 2024poster

We present DiffuScene for indoor 3D scene synthesis based on a novel scene configuration denoising diffusion model. It generates 3D instance properties stored in an unordered object set and retrieves the most similar geometry for each object configuration which is characterized as a concatenation of…

Cited by 99SourcePDFScholar
2024

FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models

CVPR 2024poster

We introduce FaceTalk a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive detailed nature of human heads including hair ears and finer-scale eye movements we propose to couple speech signal…

2024

Generating Human Interaction Motions in Scenes with Text Control

ECCV 2024poster

"We present , a text-controlled scene-aware motion generation method based on denoising diffusion models. Previous text-to-motion methods focus on characters in isolation without considering scenes due to the limited availability of datasets that include motion, text descriptions, and interactive sc…

Cited by 40SourcePDFScholar
2024

Human Hair Reconstruction with Strand-Aligned 3D Gaussians

ECCV 2024poster

"We introduce a new hair modeling method that uses a dual representation of classical hair strands and 3D Gaussians to produce accurate and realistic strand-based reconstructions from multi-view data. In contrast to recent approaches that leverage unstructured Gaussians to model human avatars, our m…

2024

SCULPT: Shape-Conditioned Unpaired Learning of Pose-dependent Clothed and Textured Human Meshes

CVPR 2024poster

We present SCULPT a novel 3D generative model for clothed and textured 3D meshes of humans. Specifically we devise a deep neural network that learns to represent the geometry and appearance distribution of clothed human bodies. Training such a model is challenging as datasets of textured 3D meshes f…

Cited by 5SourcePDFScholar
2024

Stable Video Portraits

ECCV 2024poster

"Rapid advances in the field of generative AI and text-to-image methods in particular have transformed the way we interact with and perceive computer-generated imagery today. In parallel, much progress has been made in 3D face reconstruction, using 3D Morphable Models (3DMM). In this paper, we prese…

Cited by 1SourcePDFScholar
2024

Synthesizing Environment-Specific People in Photographs

ECCV 2024poster

"We present ESP, a novel method for context-aware full-body generation, that enables photo-realistic synthesis and inpainting of people wearing clothing that is semantically appropriate for the scene depicted in an input photograph. ESP is conditioned on a 2D pose and contextual cues that are extrac…

Cited by 0SourcePDFScholar
2024

Text-Conditioned Generative Model of 3D Strand-based Human Hairstyles

CVPR 2024poster

We present HAAR a new strand-based generative model for 3D human hairstyles. Specifically based on textual inputs HAAR produces 3D hairstyles that could be used as production-level assets in modern computer graphics engines. Current AI-based generative models take advantage of powerful 2D priors to…

Cited by 3SourcePDFScholar
2023

CaPhy: Capturing Physical Properties for Animatable Human Avatars

ICCV 2023poster

We present CaPhy, a novel method for reconstructing animatable human avatars with realistic dynamic properties for clothing. Specifically, we aim for capturing the geometric and physical properties of the clothing from real observations. This allows us to apply novel poses to the human avatar with p…

Cited by 15PDFScholar
2023

High-Res Facial Appearance Capture From Polarized Smartphone Images

CVPR 2023poster

We propose a novel method for high-quality facial texture reconstruction from RGB images using a novel capturing routine based on a single smartphone which we equip with an inexpensive polarization foil. Specifically, we turn the flashlight into a polarized light source and add a polarization filter…

Cited by 17SourcePDFScholar
2023

Imitator: Personalized Speech-driven 3D Facial Animation

ICCV 2023poster

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input audio without considering the identity-specific speaking st…

Cited by 60PDFcodeScholar
2023

MIME: Human-Aware 3D Scene Generation

CVPR 2023poster

Generating realistic 3D worlds occupied by moving humans has many applications in games, architecture, and synthetic data creation. But generating such scenes is expensive and labor intensive. Recent work generates human poses and motions given a 3D scene. Here, we take the opposite approach and gen…

2022

Human-Aware Object Placement for Visual Environment Reconstruction

CVPR 2022poster

Humans are in constant contact with the world as they move through it and interact with it. This contact is a vital source of information for understanding 3D humans, 3D scenes, and the interactions between them. In fact, we demonstrate that these human-scene interactions (HSIs) can be leveraged to…

Cited by 70PDFcodeScholar
2022

Neural Head Avatars From Monocular RGB Videos

CVPR 2022poster

We present Neural Head Avatars, a novel neural representation that explicitly models the surface geometry and appearance of an animatable human avatar that can be used for teleconferencing in AR/VR or other applications in the movie or games industry that rely on a digital human. Our representation…

Cited by 223PDFScholar
2022

Neural RGB-D Surface Reconstruction

CVPR 2022poster

Obtaining high-quality 3D reconstructions of room-scale scenes is of paramount importance for upcoming applications in AR or VR. These range from mixed reality applications for teleconferencing, virtual measuring, virtual room planing, to robotic applications. While current volume-based view synthes…

Cited by 384PDFcodeScholar
2022

Texturify: Generating Textures on 3D Shape Surfaces

ECCV 2022poster

"Texture cues on 3D objects are key to compelling visual representations, with the possibility to create high visual fidelity with inherent spatial consistency across different views. Since the availability of textured 3D shapes remains very limited, learning a 3D-supervised data-driven method that…

Cited by 73SourcePDFScholar
2021

Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction

CVPR 2021poster

We present dynamic neural radiance fields for modeling the appearance and dynamics of a human face. Digitally modeling and reconstructing a talking human is a key building-block for a variety of applications. Especially, for telepresence applications in AR or VR, a faithful reproduction of the appea…

Cited by 638PDFScholar
2021

ID-Reveal: Identity-Aware DeepFake Video Detection

ICCV 2021poster

A major challenge in DeepFake forgery detection is that state-of-the-art algorithms are mostly trained to detect a specific fake method. As a result, these approaches show poor generalization across different types of facial manipulations, e.g., from face swapping to facial reenactment. To this end,…

Cited by 219PDFcodeScholar
2021

NPMs: Neural Parametric Models for 3D Deformable Shapes

ICCV 2021poster

Parametric 3D models have enabled a wide variety of tasks in computer graphics and vision, such as modeling human bodies, faces, and hands. However, the construction of these parametric models is often tedious, as it requires heavy manual tweaking, and they struggle to represent additional complexit…

Cited by 118PDFcodeScholar
2021

Neural Deformation Graphs for Globally-Consistent Non-Rigid Reconstruction

CVPR 2021poster

We introduce Neural Deformation Graphs for globally-consistent deformation tracking and 3D reconstruction of non-rigid objects. Specifically, we implicitly model a deformation graph via a deep neural network. This neural deformation graph does not rely on any object-specific structure and, thus, can…

Cited by 83PDFcodeScholar
2021

RetrievalFuse: Neural 3D Scene Reconstruction With a Database

ICCV 2021poster

3D reconstruction of large scenes is a challenging problem due to the high-complexity nature of the solution space, in particular for generative neural networks. In contrast to traditional generative learned models which encode the full generative process into a neural network and can struggle with…

Cited by 38PDFcodeScholar
2021

SPSG: Self-Supervised Photometric Scene Generation From RGB-D Scans

CVPR 2021poster

We present SPSG, a novel approach to generate high-quality, colored 3D models of scenes from RGB-D scan observations by learning to infer unobserved scene geometry and color in a self-supervised fashion. Our self-supervised approach learns to jointly inpaint geometry and color by correlating an inco…

Cited by 42PDFcodeScholar
2021

TransformerFusion: Monocular RGB Scene Reconstruction using Transformers

NeurIPS 2021poster

We introduce TransformerFusion, a transformer-based 3D scene reconstruction approach. From an input monocular RGB video, the video frames are processed by a transformer network that fuses the observations into a volumetric feature grid representing the scene; this feature grid is then decoded into a…

Cited by 158SourcePDFScholar
2020

Adversarial Texture Optimization From RGB-D Scans

CVPR 2020poster

Realistic color texture generation is an important step in RGB-D surface reconstruction, but remains challenging in practice due to inaccuracies in reconstructed geometry, misaligned camera poses, and view-dependent imaging artifacts. In this work, we present a novel approach for color texture gener…

Cited by 59PDFcodeScholar
2020

Image-guided Neural Object Rendering

ICLR 2020poster

We propose a learned image-guided rendering technique that combines the benefits of image-based rendering and GAN-based image synthesis. The goal of our method is to generate photo-realistic re-renderings of reconstructed objects for virtual and augmented reality applications (e.g., virtual showroom…

Cited by 71SourceScholar
2020

Neural Non-Rigid Tracking

NeurIPS 2020poster

We introduce a novel, end-to-end learnable, differentiable non-rigid tracker that enables state-of-the-art non-rigid reconstruction by a learned robust optimization. Given two input RGB-D frames of a non-rigidly moving object, we employ a convolutional neural network to predict dense correspondences…

2020

Neural Voice Puppetry: Audio-driven Facial Reenactment

ECCV 2020poster

We present Neural Voice Puppetry, a novel approach for audio-driven facial video synthesis. Given an audio sequence of a source person or digital assistant, we generate a photo-realistic output video of a target person that is in sync with the audio of the source input. This audio-driven facial reen…

2019

DeepVoxels: Learning Persistent 3D Feature Embeddings

CVPR 2019oral

In this work, we address the lack of 3D understanding of generative neural networks by introducing a persistent 3D feature embedding for view synthesis. To this end, we propose DeepVoxels, a learned representation that encodes the view-dependent appearance of a 3D scene without having to explicitly…

Cited by 725PDFScholar
2019

FaceForensics++: Learning to Detect Manipulated Facial Images

ICCV 2019poster

The rapid progress in synthetic image generation and manipulation has now come to a point where it raises significant concerns for the implications towards society. At best, this leads to a loss of trust in digital content, but could potentially cause further harm by spreading false information or f…

Cited by 2929PDFcodeScholar
2018

InverseFaceNet: Deep Monocular Inverse Face Rendering

CVPR 2018poster

We introduce InverseFaceNet, a deep convolutional inverse rendering framework for faces that jointly estimates facial pose, shape, expression, reflectance and illumination from a single input image. By estimating all parameters from just a single image, advanced editing possibilities on a single fac…

Cited by 76SourcePDFScholar
2016

Face2Face: Real-Time Face Capture and Reenactment of RGB Videos

CVPR 2016oral

We present a novel approach for real-time facial reenactment of a monocular target video sequence (e.g., Youtube video). The source sequence is also a monocular video stream, captured live with a commodity webcam. Our goal is to animate the facial expressions of the target video by a source actor an…

Cited by 2654PDFScholar