← Search

Otmar Hilliges

80 accepted papers

2025

Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models

ICLR 2025poster

Classifier-free guidance (CFG) is crucial for improving both generation quality and alignment between the input condition and final output in diffusion models. While a high guidance scale is generally required to enhance these aspects, it also causes oversaturation and unrealistic artifacts. In this…

Cited by 4SourcePDFScholar
2025

Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction

ICLR 2025poster

Understanding drivers’ decision-making is crucial for road safety. Although predicting the ego-vehicle’s path is valuable for driver-assistance systems, existing methods mainly focus on external factors like other vehicles’ motions, often neglecting the driver’s attention and intent. To address this…

2025

No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion Models

ICLR 2025poster

Classifier-free guidance (CFG) has become the standard method for enhancing the quality of conditional diffusion models. However, employing CFG requires either training an unconditional model alongside the main diffusion model or modifying the training procedure by periodically inserting a null cond…

Cited by 6SourcePDFScholar
2024

4D-DRESS: A 4D Dataset of Real-World Human Clothing With Semantic Annotations

CVPR 2024highlight

The studies of human clothing for digital avatars have predominantly relied on synthetic datasets. While easy to collect synthetic data often fall short in realism and fail to capture authentic clothing dynamics. Addressing this gap we introduce 4D-DRESS the first real-world 4D dataset advancing hum…

2024

A Unified Approach for Text- and Image-guided 4D Scene Generation

CVPR 2024poster

Large-scale diffusion generative models are greatly simplifying image video and 3D asset creation from user provided text prompts and images. However the challenging problem of text-to-4D dynamic 3D scene generation with diffusion guidance remains largely unexplored. We propose Dream-in-4D which fea…

Cited by 51SourcePDFScholar
2024

AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view Videos

ECCV 2024poster

"Despite progress in human motion capture, existing multi-view methods often face challenges in estimating the 3D pose and shape of multiple closely interacting people. This difficulty arises from reliance on accurate 2D joint estimations, which are hard to obtain due to occlusions and body contact…

Cited by 2SourcePDFScholar
2024

Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects

ECCV 2024poster

"We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic understanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion generation. Accurately reconstructing such interactions in is c…

2024

CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed Sampling

ICLR 2024spotlight

While conditional diffusion models are known to have good coverage of the data distribution, they still face limitations in output diversity, particularly when sampled with a high classifier-free guidance scale for optimal image quality or when trained on small datasets. We attribute this problem to…

Cited by 38SourcePDFScholar
2024

GraspXL: Generating Grasping Motions for Diverse Objects at Scale

ECCV 2024poster

"Human hands possess the dexterity to interact with diverse objects such as grasping specific parts of the objects and/or approaching them from desired directions. More importantly, humans can grasp objects of any shape without object-specific skills. Recent works synthesize grasping motions followi…

Cited by 27SourcePDFScholar
2024

HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video

CVPR 2024highlight

Since humans interact with diverse objects every day the holistic 3D capture of these interactions is important to understand and model human behaviour. However most existing methods for hand-object reconstruction from RGB either assume pre-scanned object templates or heavily rely on limited 3D hand…

2024

HSR: Holistic 3D Human-Scene Reconstruction from Monocular Videos

ECCV 2024poster

"An overarching goal for computer-aided perception systems is the holistic understanding of the human-centric 3D world, including faithful reconstructions of humans, scenes, and their global spatial relationships. While recent progress in monocular 3D reconstruction has been made for footage of eith…

Cited by 3SourcePDFScholar
2024

Human Hair Reconstruction with Strand-Aligned 3D Gaussians

ECCV 2024poster

"We introduce a new hair modeling method that uses a dual representation of classical hair strands and 3D Gaussians to produce accurate and realistic strand-based reconstructions from multi-view data. In contrast to recent approaches that leverage unstructured Gaussians to model human avatars, our m…

2024

LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models

NeurIPS 2024poster

Advances in latent diffusion models (LDMs) have revolutionized high-resolution image generation, but the design space of the autoencoder that is central to these systems remains underexplored. In this paper, we introduce LiteVAE, a new autoencoder design for LDMs, which leverages the 2D discrete wav…

Cited by 10SourcePDFScholar
2024

MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild

CVPR 2024poster

We present MultiPly a novel framework to reconstruct multiple people in 3D from monocular in-the-wild videos. Reconstructing multiple individuals moving and interacting naturally from monocular in-the-wild videos poses a challenging task. Addressing it necessitates precise pixel-level disentanglemen…

Cited by 9SourcePDFScholar
2024

PALM: Predicting Actions through Language Models

ECCV 2024poster

"Understanding human activity is a crucial yet intricate task in egocentric vision, a field that focuses on capturing visual perspectives from the camera wearer’s viewpoint. Traditional methods heavily rely on representation learning that is trained on a large amount of video data. However, a major…

2024

ReLoo: Reconstructing Humans Dressed in Loose Garments from Monocular Video in the Wild

ECCV 2024poster

"While previous years have seen great progress in the 3D reconstruction of humans from monocular videos, few of the state-of-the-art methods are able to handle loose garments that exhibit large non-rigid surface deformations during articulation. This limits the application of such methods to humans…

Cited by 7SourcePDFScholar
2024

SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion

CVPR 2024poster

A long-standing goal of 3D human reconstruction is to create lifelike and fully detailed 3D humans from single-view images. The main challenge lies in inferring unknown body shapes appearances and clothing details in areas not visible in the images. To address this we propose SiTH a novel pipeline t…

2024

Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation

CVPR 2024poster

We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects coined "action context". We propose TransFusion a multimodal transformer-based architecture for short-term object interaction anticipati…

Cited by 5SourcePDFScholar
2024

SynH2R: Synthesizing Hand-Object Motions for Learning Human-to-Robot Handovers

ICRA 2024poster

Vision-based human-to-robot handover is an important and challenging task in human-robot interaction. Recent work has attempted to train robot policies by interacting with dynamic virtual humans in simulated environments, where the policies can later be transferred to the real world. However, a majo…

Cited by 20SourceScholar
2024

Text-Conditioned Generative Model of 3D Strand-based Human Hairstyles

CVPR 2024poster

We present HAAR a new strand-based generative model for 3D human hairstyles. Specifically based on textual inputs HAAR produces 3D hairstyles that could be used as production-level assets in modern computer graphics engines. Current AI-based generative models take advantage of powerful 2D priors to…

Cited by 3SourcePDFScholar
2024

WANDR: Intention-guided Human Motion Generation

CVPR 2024poster

Synthesizing natural human motions that enable a 3D human avatar to walk and reach for arbitrary goals in 3D space remains an unsolved problem with many applications. Existing methods (data-driven or using reinforcement learning) are limited in terms of generalization and motion naturalness. A prima…

Cited by 12SourcePDFScholar
2024

WorldPose: A World Cup Dataset for Global 3D Human Pose Estimation

ECCV 2024poster

"We present , a novel dataset for advancing research in multi-person global pose estimation in the wild, featuring footage from the 2022 FIFA World Cup. While previous datasets have primarily focused on local poses, often limited to a single person or in constrained, indoor settings, the infrastruct…

Cited by 5SourcePDFScholar
2023

AG3D: Learning to Generate 3D Avatars from 2D Image Collections

ICCV 2023poster

While progress in 2D generative models of human appearance has been rapid, many applications require 3D avatars that can be animated and rendered. Unfortunately, most existing methods for learning generative models of 3D humans with diverse shape and appearance require 3D training data, which is lim…

Cited by 60PDFScholar
2023

ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

CVPR 2023poster

Humans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because there exist no datasets with ground-truth 3D annotations for…

2023

EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild

ICCV 2023poster

We present EMDB, the Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. EMDB is a novel dataset that contains high-quality 3D SMPL pose and shape parameters with global body and camera trajectories for in-the-wild videos. We use body-worn, wireless electromagnetic (EM) sensors a…

Cited by 52PDFcodeScholar
2023

Efficient Learning of High Level Plans from Play

ICRA 2023poster

Real-world robotic manipulation tasks remain an elusive challenge, since they involve both fine-grained environment interaction, as well as the ability to plan for long-horizon goals. Although deep reinforcement learning (RL) methods have shown encouraging results when planning end-to-end in high-di…

Cited by 5SourceScholar
2023

GazeNeRF: 3D-Aware Gaze Redirection With Neural Radiance Fields

CVPR 2023poster

We propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and eye balls are separate 3D structures that move in a coordinated…

2023

HARP: Personalized Hand Reconstruction From a Monocular RGB Video

CVPR 2023poster

We present HARP (HAnd Reconstruction and Personalization), a personalized hand avatar creation approach that takes a short monocular RGB video of a human hand as input and reconstructs a faithful hand avatar exhibiting a high-fidelity appearance and geometry. In contrast to the major trend of neural…

Cited by 31SourcePDFScholar
2023

HOOD: Hierarchical Graphs for Generalized Modelling of Clothing Dynamics

CVPR 2023poster

We propose a method that leverages graph neural networks, multi-level message passing, and unsupervised training to enable real-time prediction of realistic clothing dynamics. Whereas existing methods based on linear blend skinning must be trained for specific garments, our method is agnostic to bod…

2023

Hi4D: 4D Instance Segmentation of Close Human Interaction

CVPR 2023poster

We propose Hi4D, a method and dataset for the auto analysis of physically close human-human interaction under prolonged contact. Robustly disentangling several in-contact subjects is a challenging task due to occlusions and complex shapes. Hence, existing multi-view systems typically fuse 3D surface…

2023

Human from Blur: Human Pose Tracking from Blurry Images

ICCV 2023poster

We propose a method to estimate 3D human poses from substantially blurred images. The key idea is to tackle the inverse problem of image deblurring by modeling the forward problem with a 3D human model, a texture map, and a sequence of poses to describe human motion. The blurring process is then mod…

Cited by 3PDFScholar
2023

InstantAvatar: Learning Avatars From Monocular Video in 60 Seconds

CVPR 2023poster

In this paper, we take one step further towards real-world applicability of monocular neural avatar reconstruction by contributing InstantAvatar, a system that can reconstruct human avatars from a monocular video within seconds, and these avatars can be animated and rendered at an interactive rate.…

Cited by 121SourcePDFScholar
2023

Learning Human-to-Robot Handovers From Point Clouds

CVPR 2023highlight

We propose the first framework to learn control policies for vision-based human-to-robot handovers, a critical task for human-robot interaction. While research in Embodied AI has made significant progress in training robot agents in simulated environments, interacting with humans remains challenging…

Cited by 50SourcePDFScholar
2023

PointAvatar: Deformable Point-Based Head Avatars From Videos

CVPR 2023poster

The ability to create realistic animatable and relightable head avatars from casual video sequences would open up wide ranging applications in communication and entertainment. Current methods either build on explicit 3D morphable meshes (3DMM) or exploit neural implicit representations. The former a…

2023

Preface: A Data-driven Volumetric Prior for Few-shot Ultra High-resolution Face Synthesis

ICCV 2023poster

NeRFs have enabled highly realistic synthesis of human faces including complex appearance and reflectance effects of hair and skin. These methods typically require a large number of multi-view input images, making the process hardware intensive and cumbersome, limiting applicability to unconstrained…

Cited by 22PDFScholar
2023

Vid2Avatar: 3D Avatar Reconstruction From Videos in the Wild via Self-Supervised Scene Decomposition

CVPR 2023poster

We present Vid2Avatar, a method to learn human avatars from monocular in-the-wild videos. Reconstructing humans that move naturally from monocular in-the-wild videos is difficult. Solving it requires accurately separating humans from arbitrary backgrounds. Moreover, it requires reconstructing detail…

2023

X-Avatar: Expressive Human Avatars

CVPR 2023poster

We present X-Avatar, a novel avatar model that captures the full expressiveness of digital humans to bring about life-like experiences in telepresence, AR/VR and beyond. Our method models bodies, hands, facial expressions and appearance in a holistic fashion and can be learned from either full 3D sc…

2022

D-Grasp: Physically Plausible Dynamic Grasp Synthesis for Hand-Object Interactions

CVPR 2022poster

We introduce the dynamic grasp synthesis task: given an object with a known 6D pose and a grasp reference, our goal is to generate motions that move the object to a target 6D pose. This is challenging, because it requires reasoning about the complex articulation of the human hand and the intricate p…

Cited by 111PDFcodeScholar
2022

I M Avatar: Implicit Morphable Head Avatars From Videos

CVPR 2022oral

Traditional 3D morphable face models (3DMMs) provide fine-grained control over expression but cannot easily capture geometric and appearance details. Neural volumetric representations approach photorealism but are hard to animate and do not generalize well to unseen expressions. To tackle this probl…

Cited by 254PDFcodeScholar
2022

LiP-Flow: Learning Inference-Time Priors for Codec Avatars via Normalizing Flows in Latent Space

ECCV 2022poster

"Neural face avatars that are trained from multi-view data captured in camera domes can produce photo-realistic 3D reconstructions. However, at inference time, they must be driven by limited inputs such as partial views recorded by headset-mounted cameras or a front-facing camera, and sparse facial…

Cited by 1SourcePDFScholar
2022

PINA: Learning a Personalized Implicit Neural Avatar From a Single RGB-D Video Sequence

CVPR 2022poster

We present a novel method to learn Personalized Implicit Neural Avatars (PINA) from a short RGB-D sequence. This allows non-expert users to create a detailed and personalized virtual copy of themselves, which can be animated with realistic clothing deformations. PINA does not require complete scans,…

Cited by 73PDFScholar
2022

SAGA: Stochastic Whole-Body Grasping with Contact

ECCV 2022poster

"The synthesis of human grasping has numerous applications including AR/VR, video games and robotics. While methods have been proposed to generate realistic hand-object interaction for object grasping and manipulation, these typically only consider interacting hand alone. Our goal is to synthesize w…

2022

gDNA: Towards Generative Detailed Neural Avatars

CVPR 2022poster

To make 3D human avatars widely available, we must be able to generate a variety of 3D virtual humans with varied identities and shapes in arbitrary poses. This task is challenging due to the diversity of clothed body shapes, their complex articulations, and the resulting rich, yet stochastic geomet…

Cited by 84PDFScholar
2021

EM-POSE: 3D Human Pose Estimation From Sparse Electromagnetic Trackers

ICCV 2021poster

Fully immersive experiences in AR/VR depend on reconstructing the full body pose of the user without restricting their motion. In this paper we study the use of body-worn electromagnetic (EM) field-based sensing for the task of 3D human pose reconstruction. To this end, we present a method to estima…

Cited by 40PDFcodeScholar
2021

Improved Learning of Robot Manipulation Tasks Via Tactile Intrinsic Motivation

RA-L 2021

In this letter we address the challenge of exploration in deep reinforcement learning for robotic manipulation tasks. In sparse goal settings, an agent does not receive any positive feedback until randomly achieving the goal, which becomes infeasible for longer control sequences. Inspired by touch-b

Cited by 31SourceScholar
2021

Learning Functionally Decomposed Hierarchies for Continuous Control Tasks With Path Planning

RA-L 2021

We present HiDe, a novel hierarchical reinforcement learning architecture that successfully solves long horizon control tasks and generalizes to unseen test scenarios. Functional decomposition between planning and low-level control is achieved by explicitly separating the state-action spaces across

Cited by 31SourceScholar
2021

PARE: Part Attention Regressor for 3D Human Body Estimation

ICCV 2021poster

Despite significant progress, we show that state of the art 3D human pose and shape estimation methods remain sensitive to partial occlusion and can produce dramatically wrong predictions although much of the body is observable. To address this, we introduce a soft attention mechanism, called the Pa…

Cited by 480PDFcodeScholar
2021

SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes

ICCV 2021poster

Neural implicit surface representations have emerged as a promising paradigm to capture 3D shapes in a continuous and resolution-independent manner. However, adapting them to articulated shapes is non-trivial. Existing approaches learn a backward warp field that maps deformed to canonical points. Ho…

Cited by 258PDFcodeScholar
2021

SPEC: Seeing People in the Wild With an Estimated Camera

ICCV 2021poster

Due to the lack of camera parameter information for in-the-wild images, existing 3D human pose and shape (HPS) estimation methods make several simplifying assumptions: weak-perspective projection, large constant focal length, and zero camera rotation. These assumptions often do not hold and we show,…

Cited by 160PDFcodeScholar
2021

Self-Supervised 3D Hand Pose Estimation From Monocular RGB via Contrastive Learning

ICCV 2021poster

Encouraged by the success of contrastive learning on image classification tasks, we propose a new self-supervised method for the structured regression task of 3D hand pose estimation. Contrastive learning makes use of unlabeled data for the purpose of representation learning via a loss formulation t…

Cited by 85PDFScholar
2021

VariTex: Variational Neural Face Textures

ICCV 2021poster

Deep generative models can synthesize photorealistic images of human faces with novel identities.However, a key challenge to the wide applicability of such techniques is to provide independent control over semantically meaningful parameters: appearance, head pose, face shape, and facial expressions.…

Cited by 43PDFcodeScholar
2020

Category Level Object Pose Estimation via Neural Analysis-by-Synthesis

ECCV 2020poster

Many object pose estimation algorithms rely on the analysis-by-synthesis framework which requires explicit representations of individual object instances. In this paper we combine a gradient-based fitting procedure with a parametric neural image synthesis module that is capable of implicitly represe…

Cited by 143SourcePDFScholar
2020

CoSE: Compositional Stroke Embeddings

NeurIPS 2020poster

We present a generative model for stroke-based drawing tasks which is able to model complex free-form structures. While previous approaches rely on sequence-based models for drawings of basic objects or handwritten text, we propose a model that treats drawings as a collection of strokes that can be…

2020

ETH-XGaze: A Large Scale Dataset for Gaze Estimation under Extreme Head Pose and Gaze Variation

ECCV 2020poster

Gaze estimation is a fundamental task in many applications of computer vision, human computer interaction and robotics. Many state-of-the-art methods are trained and tested on custom datasets, making comparison across methods challenging. Furthermore, existing gaze estimation datasets have limited h…

2020

Learning to Assemble: Estimating 6D Poses for Robotic Object-Object Manipulation

RA-L 2020

In this letter we propose a robotic vision task with the goal of enabling robots to execute complex assembly tasks in unstructured environments using a camera as the primary sensing device. We formulate the task as an instance of 6D pose estimation of template geometries, to which manipulation objec

Cited by 35SourceScholar
2020

Self-Learning Transformations for Improving Gaze and Head Redirection

NeurIPS 2020poster

Many computer vision tasks rely on labeled data. Rapid progress in generative modeling has led to the ability to synthesize photorealistic images. However, controlling specific aspects of the generation process such that the data can be used for supervision of downstream tasks remains challenging. I…

2020

Weakly Supervised 3D Hand Pose Estimation via Biomechanical Constraints

ECCV 2020poster

Estimating 3D hand pose from 2D images is a difficult, inverse problem due to the inherent scale and depth ambiguities. Current state-of-the-art methods train fully supervised deep neural networks with 3D ground-truth data. However, acquiring 3D annotations is expensive, typically requiring calibrat…

Cited by 185SourcePDFScholar
2019

Demonstration-Guided Deep Reinforcement Learning of Control Policies for Dexterous Human-Robot Interaction

ICRA 2019poster

In this paper, we propose a method for training control policies for human-robot interactions such as handshakes or hand claps via Deep Reinforcement Learning. The policy controls a humanoid Shadow Dexterous Hand, attached to a robot arm. We propose a parameterizable multi-objective reward function…

Cited by 53SourceScholar
2019

Few-Shot Adaptive Gaze Estimation

ICCV 2019oral

Inter-personal anatomical differences limit the accuracy of person-independent gaze estimation networks. Yet there is a need to lower gaze errors further to enable applications requiring higher quality. Further gains can be achieved by personalizing gaze networks, ideally with few calibration sample…

Cited by 248PDFcodeScholar
2019

Photo-Realistic Monocular Gaze Redirection Using Generative Adversarial Networks

ICCV 2019poster

Gaze redirection is the task of changing the gaze to a desired direction for a given monocular eye patch image. Many applications such as videoconferencing, films, games, and generation of training data for gaze estimation require redirecting the gaze, without distorting the appearance of the area s…

Cited by 63PDFcodeScholar
2019

Video-based Prediction of Hand-grasp Preshaping with Application to Prosthesis Control

ICRA 2019poster

Among the currently available grasp-type selection techniques for hand prostheses, there is a distinct lack of intuitive, robust, low-latency solutions. In this paper we investigate the use of a portable, forearm-mounted, video-based technique for the prediction of hand-grasp preshaping for arbitrar…

Cited by 32SourceScholar
2018

Deformation Capture via Self-Sensing Capacitive Arrays (Video)

IROS 2018poster

In this video we present soft self-sensing capacitive arrays and demonstrate their use in capturing dense surface deformations without requiring line of sight. The capacitive arrays are made of two electrode patterns embedded into a single silicone compound. The overlaps of the electrode strip patte…

Cited by 1SourceScholar
2018

Learn-to-Score: Efficient 3D Scene Exploration by Predicting View Utility

ECCV 2018poster

Camera equipped drones are nowadays being used to explore large scenes and reconstruct detailed 3D maps. When free space in the scene is approximately known, an offline planner can generate optimal plans to efficiently explore the scene. However, for exploring unknown scenes, the planner must predic…

Cited by 62SourcePDFScholar
2018

Sample Efficient Learning of Path Following and Obstacle Avoidance Behavior for Quadrotors

RA-L 2018

In this letter, we propose an algorithm for the training of neural network control policies for quadrotors. The learned control policy computes control commands directly from sensor inputs and is, hence, computationally efficient. An imitation learning algorithm produces a policy that reproduces the

Cited by 20SourceScholar
2017

Real-Time Motion Planning for Aerial Videography With Real-Time With Dynamic Obstacle Avoidance and Viewpoint Optimization

RA-L 2017

We propose a method for real-time trajectory generation with applications in aerial videography. Taking framing objectives, such as position of targets in the image plane, as input, our method solves for robot trajectories and gimbal controls automatically and adapts plans in real time due to change

Cited by 83SourceScholar
2017

Thin-Slicing Network: A Deep Structured Model for Pose Estimation in Videos

CVPR 2017oral

Deep ConvNets have been shown to be effective for the task of human pose estimation from single images. However, several challenging issues arise in the video-based case such as self-occlusion, motion blur, and uncommon poses with few or no examples in the training data. Temporal information can pro…

Cited by 159PDFScholar
2016

Omni-directional person tracking on a flying robot using occlusion-robust ultra-wideband signals

IROS 2016poster

We present a tracking system based on ultra-wideband (UWB) radio tranceivers mounted on a robot and a target. In comparison to typical UWB localization systems with fixed UWB tranceivers in the environment we only require instrumentation of the target with a single UWB tranceiver. Our system works i…

Cited by 62SourceScholar
2015

Semi-direct EKF-based monocular visual-inertial odometry

IROS 2015poster

We propose a novel monocular visual inertial odometry algorithm that combines the advantages of EKF-based approaches with those of direct photometric error minimization methods. The method is based on sparse, very small patches and incorporates the minimization of photometric error directly into the…

Cited by 75SourceScholar