← Search

Carsten Stoll

6 accepted papers

2026

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

CVPR 2026

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and

Cited by 0SourcecodeScholar
2024

PhISANet: Phonetically Informed Speech Animation Network

ICASSP 2024accepted

Realistic animation is crucial for immersive and seamless human-avatar interactions as digital avatars become more prevalent. This work presents PhISANet, an encoder-decoder model that realistically animates the face and tongue solely from speech. PhISANet leverages neural audio representations trai…

Cited by 0SourceScholar
2022

Speech Driven Tongue Animation

CVPR 2022poster

Advances in speech driven animation techniques allow the creation of convincing animations for virtual characters solely from audio data. Many existing approaches focus on facial and lip motion and they often do not provide realistic animation of the inner mouth. This paper addresses the problem of…

Cited by 21PDFScholar
2021

ANR: Articulated Neural Rendering for Virtual Avatars

CVPR 2021poster

Deferred Neural Rendering (DNR) uses a three-step pipeline to translate a mesh representation into an RGB image. The combination of a traditional rendering stack with neural networks hits a sweet spot in terms of computational complexity and realism of the resulting images. Using skinned meshes for…

Cited by 70PDFScholar
2020

PatchNets: Patch-Based Generalizable Deep Implicit 3D Shape Representations

ECCV 2020poster

Implicit surface representation combined with deep learning has led to impressive models which can represent detailed shapes of objects. Implicit surface representations, such as signed-distance functions, allow to represent shapes of arbitrary topologies. Since a continous function is learned, the…

Cited by 116SourcePDFScholar
2020

TexMesh: Reconstructing Detailed Human Texture and Geometry from RGB-D Video

ECCV 2020poster

We present TexMesh, a novel approach to reconstruct detailed human meshes with high-resolution full-body texture from RGB-D video. TexMesh enables high quality free-viewpoint rendering of humans. Given the RGB frames, the captured environment map, and the coarse per-frame human mesh from RGB-D track…

Cited by 53SourcePDFScholar