← Search

Manuel Kaufmann

13 accepted papers

2026

RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos

CVPR 2026

Reconstructing people, objects, and their interactions in 3D is a long-standing and fundamental goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each other, and camera and object motion entangle

Cited by 0SourcecodeScholar
2025

ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular Videos

CVPR 2025poster

Creating a photorealistic scene and human reconstruction from a single monocular in-the-wild video figures prominently in the perception of a human-centric 3D world. Recent neural rendering advances have enabled holistic human-scene reconstruction but require pre-calibrated camera and human poses, a…

2025

PHD: Personalized 3D Human Body Fitting with Point Diffusion

ICCV 2025poster

We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these met…

2024

HSR: Holistic 3D Human-Scene Reconstruction from Monocular Videos

ECCV 2024poster

"An overarching goal for computer-aided perception systems is the holistic understanding of the human-centric 3D world, including faithful reconstructions of humans, scenes, and their global spatial relationships. While recent progress in monocular 3D reconstruction has been made for footage of eith…

Cited by 3SourcePDFScholar
2024

MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild

CVPR 2024poster

We present MultiPly a novel framework to reconstruct multiple people in 3D from monocular in-the-wild videos. Reconstructing multiple individuals moving and interacting naturally from monocular in-the-wild videos poses a challenging task. Addressing it necessitates precise pixel-level disentanglemen…

Cited by 9SourcePDFScholar
2024

ReLoo: Reconstructing Humans Dressed in Loose Garments from Monocular Video in the Wild

ECCV 2024poster

"While previous years have seen great progress in the 3D reconstruction of humans from monocular videos, few of the state-of-the-art methods are able to handle loose garments that exhibit large non-rigid surface deformations during articulation. This limits the application of such methods to humans…

Cited by 7SourcePDFScholar
2024

WorldPose: A World Cup Dataset for Global 3D Human Pose Estimation

ECCV 2024poster

"We present , a novel dataset for advancing research in multi-person global pose estimation in the wild, featuring footage from the 2022 FIFA World Cup. While previous datasets have primarily focused on local poses, often limited to a single person or in constrained, indoor settings, the infrastruct…

Cited by 5SourcePDFScholar
2023

ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

CVPR 2023poster

Humans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because there exist no datasets with ground-truth 3D annotations for…

2023

EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild

ICCV 2023poster

We present EMDB, the Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. EMDB is a novel dataset that contains high-quality 3D SMPL pose and shape parameters with global body and camera trajectories for in-the-wild videos. We use body-worn, wireless electromagnetic (EM) sensors a…

Cited by 52PDFcodeScholar
2023

Hi4D: 4D Instance Segmentation of Close Human Interaction

CVPR 2023poster

We propose Hi4D, a method and dataset for the auto analysis of physically close human-human interaction under prolonged contact. Robustly disentangling several in-contact subjects is a challenging task due to occlusions and complex shapes. Hence, existing multi-view systems typically fuse 3D surface…

2023

X-Avatar: Expressive Human Avatars

CVPR 2023poster

We present X-Avatar, a novel avatar model that captures the full expressiveness of digital humans to bring about life-like experiences in telepresence, AR/VR and beyond. Our method models bodies, hands, facial expressions and appearance in a holistic fashion and can be learned from either full 3D sc…

2021

EM-POSE: 3D Human Pose Estimation From Sparse Electromagnetic Trackers

ICCV 2021poster

Fully immersive experiences in AR/VR depend on reconstructing the full body pose of the user without restricting their motion. In this paper we study the use of body-worn electromagnetic (EM) field-based sensing for the task of 3D human pose reconstruction. To this end, we present a method to estima…

Cited by 40PDFcodeScholar