← Search

Marc Habermann

29 accepted papers

2026

$\boldsymbol{\partial^\infty}$-Grid: A Neural Differential Equation Solver with Differentiable Feature Grids

ICLR 2026poster

We present a novel differentiable grid-based representation for efficiently solving differential equations (DEs). Widely used architectures for neural solvers, such as sinusoidal neural networks, are coordinate-based MLPs that are, both, computationally intensive and slow to train. Although grid-bas…

Cited by 0SourcecodeScholar
2026

OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control

CVPR 2026

We introduce OLATverse, a large-scale dataset comprising around 9M images of 765 real-world objects, captured from multiple viewpoints under a diverse set of precisely controlled lighting conditions. While recent advances in object-centric inverse rendering, novel view synthesis and relighting have

Cited by 0SourcecodeScholar
2026

RelightAnyone: A Generalized Relightable 3D Gaussian Head Model

CVPR 2026

3D Gaussian Splatting (3DGS) has become a standard approach to reconstruct and render photorealistic 3D head avatars. A major challenge is to relight the avatars to match any scene illumination. For high quality relighting, existing methods require subjects to be captured under complex time-multiple

Cited by 0SourceScholar
2026

Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views

CVPR 2026

We present _Relightable Holoported Characters_ (RHC), a novel person-specific method for free-view rendering and relighting of full-body and highly dynamic humans solely observed from sparse-view RGB videos at inference. In contrast to classical one-light-at-a-time (OLAT)-based human relighting, our

Cited by 0SourceScholar
2025

BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects

CVPR 2025poster

We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-base…

Cited by 2SourcePDFScholar
2025

EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild

CVPR 2025poster

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task.Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view reconstruction faces significant challenges due to inherent ambiguities an…

2025

FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video

CVPR 2025highlight

Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurat…

Cited by 0SourcePDFScholar
2025

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

CVPR 2025poster

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform general dynamic scenes. In this paper, we present RePerformer, a novel Gaussian-based rep…

Cited by 0SourcePDFScholar
2025

Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures

CVPR 2025highlight

Real-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space to learn rendering primitives. However, they either jointly learn geometry and…

Cited by 2SourcePDFScholar
2025

Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation Fields

CVPR 2025poster

3D reconstruction of highly deformable surfaces (e.g. cloths) from monocular RGB videos is a challenging problem, and no solution provides a consistent and accurate recovery of fine-grained surface details. To account for the ill-posed nature of the setting, existing methods use deformation models w…

Cited by 0SourcePDFScholar
2024

ASH: Animatable Gaussian Splats for Efficient and Photoreal Human Rendering

CVPR 2024poster

Real-time rendering of photorealistic and controllable human avatars stands as a cornerstone in Computer Vision and Graphics. While recent advances in neural implicit rendering have unlocked unprecedented photorealism for digital avatars real-time performance has mostly been demonstrated for static…

2024

ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis

CVPR 2024poster

Gestures play a key role in human communication. Recent methods for co-speech gesture generation while managing to generate beat-aligned motions struggle generating gestures that are semantically aligned with the utterance. Compared to beat gestures that align naturally to the audio signal semantica…

Cited by 12SourcePDFScholar
2024

Holoported Characters: Real-time Free-viewpoint Rendering of Humans from Sparse RGB Cameras

CVPR 2024poster

We present the first approach to render highly realistic free-viewpoint videos of a human actor in general apparel from sparse multi-view recording to display in real-time at an unprecedented 4K resolution. At inference our method only requires four camera views of the moving actor and the respectiv…

Cited by 9SourcePDFScholar
2024

MetaCap: Meta-learning Priors from Multi-View Imagery for Sparse-view Human Performance Capture and Rendering

ECCV 2024poster

"Faithful human performance capture and free-view rendering from sparse RGB observations is a long-standing problem in Vision and Graphics. The main challenges are the lack of observations and the inherent ambiguities of the setting, e.g. occlusions and depth ambiguity. As a result, radiance fields,…

Cited by 9SourcePDFScholar
2024

NeuralClothSim: Neural Deformation Fields Meet the Thin Shell Theory

NeurIPS 2024poster

Despite existing 3D cloth simulators producing realistic results, they predominantly operate on discrete surface representations (e.g. points and meshes) with a fixed spatial resolution, which often leads to large memory consumption and resolution-dependent simulations. Moreover, back-propagating gr…

2024

Relightable Neural Actor with Intrinsic Decomposition and Pose Control

ECCV 2024poster

"Creating a controllable and relightable digital avatar from multi-view video with fixed illumination is a very challenging problem since humans are highly articulated, creating pose-dependent appearance effects, and skin as well as clothing require space-varying BRDF modeling. Existing works on cre…

Cited by 4SourcePDFScholar
2024

Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models

ECCV 2024poster

"We present Surf-D, a novel method for generating high-quality 3D shapes as Surfaces with arbitrary topologies using Diffusion models. Previous methods explored shape generation with different representations and they suffer from limited topologies and poor geometry details. To generate high-quality…

Cited by 1SourcePDFScholar
2024

VINECS: Video-based Neural Character Skinning

CVPR 2024poster

Rigging and skinning clothed human avatars is a challenging task and traditionally requires a lot of manual work and expertise. Recent methods addressing it either generalize across different characters or focus on capturing the dynamics of a single character observed under different pose configurat…

Cited by 3SourcePDFScholar
2024

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

CVPR 2024highlight

In this work we introduce Wonder3D a novel method for generating high-fidelity textured meshes from single-view images with remarkable efficiency. Recent methods based on the Score Distillation Sampling (SDS) loss methods have shown the potential to recover 3D geometry from 2D diffusion priors but t…

Cited by 414SourcePDFScholar
2023

DELIFFAS: Deformable Light Fields for Fast Avatar Synthesis

NeurIPS 2023poster

Generating controllable and photorealistic digital human avatars is a long-standing and important problem in Vision and Graphics. Recent methods have shown great progress in terms of either photorealism or inference speed while the combination of the two desired properties still remains unsolved. To…

Cited by 33SourcePDFScholar
2023

LiveHand: Real-time and Photorealistic Neural Hand Rendering

ICCV 2023poster

The human hand is the main medium through which we interact with our surroundings, making its digitization an important problem. While there are several works modeling the geometry of hands, little attention has been paid to capturing photo-realistic appearance. Moreover, for applications in extende…

Cited by 19PDFcodeScholar
2023

NeuS2: Fast Learning of Neural Implicit Surfaces for Multi-view Reconstruction

ICCV 2023poster

Recent methods for neural surface representation and rendering, for example NeuS, have demonstrated the remarkably high-quality reconstruction of static scenes. However, the training of NeuS takes an extremely long time (8 hours), which makes it almost impossible to apply them to dynamic scenes with…

Cited by 276PDFcodeScholar
2022

Neural Radiance Transfer Fields for Relightable Novel-View Synthesis with Global Illumination

ECCV 2022poster

"Given a set of images of a scene, the re-rendering of this scene from novel views and lighting conditions is an important and challenging problem in Computer Vision and Graphics. On the one hand, most existing works in Computer Vision usually impose many assumptions regarding the image formation pr…

Cited by 50SourcePDFScholar
2022

Physical Inertial Poser (PIP): Physics-Aware Real-Time Human Motion Tracking From Sparse Inertial Sensors

CVPR 2022poster

Motion capture from sparse inertial sensors has shown great potential compared to image-based approaches since occlusions do not lead to a reduced tracking quality and the recording space is not restricted to be within the viewing frustum of the camera. However, capturing the motion and global posit…

Cited by 200PDFScholar
2021

Efficient and Differentiable Shadow Computation for Inverse Problems

ICCV 2021poster

Differentiable rendering has received increasing interest in the solution of image-based inverse problems. It can benefit traditional optimization-based solutions to inverse problems, but also allows for self-supervision of learning-based approaches for which training data with ground truth annotati…

Cited by 14PDFScholar
2021

Monocular Real-Time Full Body Capture With Inter-Part Correlations

CVPR 2021poster

We present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network architecture that exploits correlations between body and hands at high computational e…

Cited by 72PDFScholar
2020

DeepCap: Monocular Human Performance Capture Using Weak Supervision

CVPR 2020oral

Human performance capture is a highly important computer vision problem with many applications in movie production and virtual/augmented reality. Many previous performance capture approaches either required expensive multi-view setups or did not recover dense space-time coherent geometry with frame-…

Cited by 266PDFScholar
2020

EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event Camera

CVPR 2020oral

The high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement, the high data bandwidth and the consequent high computation overhead. In this paper, we propose EventCap -- the first…

Cited by 124PDFScholar
2020

Monocular Real-Time Hand Shape and Motion Capture Using Multi-Modal Data

CVPR 2020poster

We present a novel method for monocular hand shape and pose estimation at unprecedented runtime performance of 100fps and at state-of-the-art accuracy. This is enabled by a new learning based architecture designed such that it can make use of all the sources of available hand training data: image da…

Cited by 257PDFcodeScholar