← Search

Thabo Beeler

23 accepted papers

2026

Physical Simulator In-the-Loop Video Generation

CVPR 2026

Recent advances in diffusion-based video generation have achieved remarkable visual realism but still struggle to obey basic physical laws such as gravity, inertia, and collision. Generated objects often move inconsistently across frames, exhibit implausible dynamics, or violate physical constraints

Cited by 0SourcecodeScholar
2026

Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views

CVPR 2026

We present _Relightable Holoported Characters_ (RHC), a novel person-specific method for free-view rendering and relighting of full-body and highly dynamic humans solely observed from sparse-view RGB videos at inference. In contrast to classical one-light-at-a-time (OLAT)-based human relighting, our

Cited by 0SourceScholar
2025

BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects

CVPR 2025poster

We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-base…

Cited by 2SourcePDFScholar
2025

Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input

CVPR 2025poster

This work focuses on tracking and understanding human motion using consumer wearable devices, such as VR/AR headsets, smart glasses, cellphones, and smartwatches. These devices provide diverse, multi-modal sensor inputs, including egocentric images, and 1-3 sparse IMU sensors in varied combinations.…

Cited by 0SourcePDFScholar
2025

FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video

CVPR 2025highlight

Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurat…

Cited by 0SourcePDFScholar
2025

GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance Modeling

CVPR 2025poster

We present GroomLight, a novel method for relightable hair appearance modeling from multi-view images. Existing hair capture methods struggle to balance photorealistic rendering with relighting capabilities. Analytical material models, while physically grounded, often fail to fully capture appearanc…

2025

Synthetic Prior for Few-Shot Drivable Head Avatar Inversion

CVPR 2025poster

We present SynShot, a novel method for the few-shot inversion of a drivable head avatar based on a synthetic prior. We tackle three major challenges. First, training a controllable 3D generative network requires a large number of diverse sequences, for which pairs of images and high-quality tracked…

Cited by 1SourcePDFScholar
2024

Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion Refinement

CVPR 2024poster

In this work we explore egocentric whole-body motion capture using a single fisheye camera which simultaneously estimates human body and hand motion. This task presents significant challenges due to three factors: the lack of high-quality datasets fisheye camera distortion and human body self-occlus…

Cited by 21SourcePDFScholar
2024

MagicMirror: Fast and High-Quality Avatar Generation with Constrained Search Space

ECCV 2024poster

"We introduce a novel framework for 3D human avatar generation and personalization, leveraging text prompts to enhance user engagement and customization. Central to our approach are key innovations aimed at overcoming the challenges in photo-realistic avatar synthesis. Firstly, we utilize a conditio…

2024

Optimizing Diffusion Noise Can Serve As Universal Motion Priors

CVPR 2024poster

We propose Diffusion Noise Optimization (DNO) a new method that effectively leverages existing motion diffusion models as motion priors for a wide range of motion-related tasks. Instead of training a task-specific diffusion model for each new task DNO operates by optimizing the diffusion latent nois…

Cited by 41SourcePDFScholar
2024

PhysAvatar: Learning the Physics of Dressed 3D Avatars from Visual Observations

ECCV 2024poster

"[width=0.9]figure/teaserv 4.pdf Figure 1: PhysAvatar is a novel framework that captures the physics of dressed 3D avatars from visual observations, enabling a wide spectrum of applications, such as (a) animation, (b) relighting, and (c) redressing, with high-fidelity rendering results."

2023

ITI-GEN: Inclusive Text-to-Image Generation

ICCV 2023oral

Text-to-image generative models often reflect the biases of the training data, leading to unequal representations of underrepresented groups. This study investigates inclusive text-to-image generative models that generate images based on human-written prompts and ensure the resulting images are unif…

Cited by 66PDFcodeScholar
2023

Learning Personalized High Quality Volumetric Head Avatars From Monocular RGB Videos

CVPR 2023poster

We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracki…

Cited by 20SourcePDFScholar
2023

Preface: A Data-driven Volumetric Prior for Few-shot Ultra High-resolution Face Synthesis

ICCV 2023poster

NeRFs have enabled highly realistic synthesis of human faces including complex appearance and reflectance effects of hair and skin. These methods typically require a large number of multi-view input images, making the process hardware intensive and cumbersome, limiting applicability to unconstrained…

Cited by 22PDFScholar
2023

Spectral Graphormer: Spectral Graph-Based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images

ICCV 2023poster

We propose a novel transformer-based framework that reconstructs two high fidelity hands from multi-view RGB images. Unlike existing hand pose estimation methods, where one typically trains a deep network to regress hand model parameters from single RGB image, we consider a more challenging problem…

Cited by 4PDFScholar
2023

Synthesizing Diverse Human Motions in 3D Indoor Scenes

ICCV 2023poster

We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on high-quality training sequences that contain captured human motions and the 3D scenes they interact with. Ho…

Cited by 66PDFcodeScholar
2022

Compositional Human-Scene Interaction Synthesis with Semantic Control

ECCV 2022poster

"Synthesizing natural interactions between virtual humans and their 3D environments is critical for numerous applications, such as computer games and AR/VR experiences. Recent methods mainly focus on modeling geometric relations between 3D environments and humans, where the high-level semantics of t…

2021

VariTex: Variational Neural Face Textures

ICCV 2021poster

Deep generative models can synthesize photorealistic images of human faces with novel identities.However, a key challenge to the wide applicability of such techniques is to provide independent control over semantically meaningful parameters: appearance, head pose, face shape, and facial expressions.…

Cited by 43PDFcodeScholar
2020

Attention-Driven Cropping for Very High Resolution Facial Landmark Detection

CVPR 2020poster

Facial landmark detection is a fundamental task for many consumer and high-end applications and is almost entirely solved by machine learning methods today. Existing datasets used to train such algorithms are primarily made up of only low resolution images, and current algorithms are limited to inpu…

Cited by 87PDFScholar
2020

ETH-XGaze: A Large Scale Dataset for Gaze Estimation under Extreme Head Pose and Gaze Variation

ECCV 2020poster

Gaze estimation is a fundamental task in many applications of computer vision, human computer interaction and robotics. Many state-of-the-art methods are trained and tested on custom datasets, making comparison across methods challenging. Furthermore, existing gaze estimation datasets have limited h…

2017

A Practical Method for Fully Automatic Intrinsic Camera Calibration Using Directionally Encoded Light

CVPR 2017spotlight

Calibrating the intrinsic properties of a camera is one of the fundamental tasks required for a variety of computer vision and image processing tasks. The precise measurement of focal length, location of the principal point as well as distortion parameters of the lens is crucial, for example, for 3D…

Cited by 7PDFcodeScholar
2015

FaceDirector: Continuous Control of Facial Performance in Video

ICCV 2015poster

We present a method to continuously blend between multiple facial performances of an actor, which can contain different facial expressions or emotional states. As an example, given sad and angry video takes of a scene, our method empowers the movie director to specify arbitrary weighted combinations…

Cited by 19PDFScholar