← Search

Gerard Pons-Moll

58 accepted papers

2026

CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

CVPR 2026

Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming, and robot learning. However, inferring 4D interactions from a single RGB view is highly challenging due to the unknown object and human information, dep

Cited by 0SourcecodeScholar
2026

CLUTCH: Contextualized Language model for Unlocking Text-Conditioned Hand motion modelling in the wild

ICLR 2026poster

Hands play a central role in daily life, yet modeling natural hand motions remains underexplored. Existing methods that tackle text-to-hand-motion generation or hand animation captioning rely on studio-captured datasets with limited actions and contexts, making them costly to scale to “in-the-wild”…

Cited by 0SourceScholar
2026

ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and Test-time Generative Adaptation

CVPR 2026

We introduce ELITE, an Efficient Gaussian head avatar synthesis from a monocular video via Learned Initialization and TEst-time generative adaptation. Prior works rely either on a 3D data prior or a 2D generative prior to compensate for missing visual cues in monocular videos. However, 3D data prior

Cited by 0SourcecodeScholar
2026

FrankenMotion: Part-level Human Motion Generation and Composition

CVPR 2026

Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptions due to the absence of fine-grained, part-level motion annotations. This limits their controllability over individual

Cited by 0SourcecodeScholar
2026

GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

CVPR 2026

Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error accumulation, or they do not explicitly leverage 3D geometry du

Cited by 0SourceScholar
2026

Human3R: Everyone Everywhere All at Once

ICLR 2026poster

We present Human3R, a unified, feed-forward framework for online 4D human-scene reconstruction, in the world frame, from casually captured monocular videos. Unlike previous approaches that rely on multi-stage pipelines, iterative contact-aware refinement between humans and scenes, and heavy dependen…

Cited by 0SourcecodeScholar
2026

MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation

CVPR 2026

We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diffusion, either on the whole latent at once or auto-regressively over multiple latents. In this paper, we study how to mak

Cited by 0SourceScholar
2026

PhysHead: Simulation-Ready Gaussian Head Avatars

CVPR 2026

Realistic digital avatars require expressive and dynamic hair motion; however, most existing head avatar methods assume rigid hair movement. These methods often fail to disentangle hair from the head, representing it as a simple outer shell and failing to capture its natural volumetric behavior. In

Cited by 0SourcecodeScholar
2025

Feat2GS: Probing Visual Foundation Models with Gaussian Splatting

CVPR 2025poster

Given that visual foundation models (VFMs) are trained on extensive datasets but often limited to 2D images, a natural question arises: how well do they understand the 3D world? With the differences in architecture and training protocols (i.e., objectives, proxy tasks), a unified framework to fairly…

2025

MVGBench: a Comprehensive Benchmark for Multi-view Generation Models

ICCV 2025poster

We propose MVGBench, a comprehensive benchmark for multi-view image generation models (MVGs) that evaluates 3D consistency in geometry and texture, image quality, and semantics (using vision language models). Recently, MVGs have been the main driving force in 3D object creation. However, existing me…

Cited by 0SourcePDFScholar
2025

TriDi: Trilateral Diffusion of 3D Humans, Objects, and Interactions

ICCV 2025poster

Modeling 3D human-object interaction (HOI) is a problem of great interest for computer vision and a key enabler for virtual and mixed-reality applications. Existing methods work in a one-way direction: some recover plausible human interactions conditioned on a 3D object; others recover the object po…

Cited by 0SourcePDFScholar
2024

GEARS: Local Geometry-aware Hand-object Interaction Synthesis

CVPR 2024poster

Generating realistic hand motion sequences in interaction with objects has gained increasing attention with the growing interest in digital humans. Prior work has illustrated the effectiveness of employing occupancy-based or distance-based virtual sensors to extract hand-object interaction features.…

Cited by 10SourcePDFScholar
2024

Human-3Diffusion: Realistic Avatar Creation via Explicit 3D Consistent Diffusion Models

NeurIPS 2024poster

Creating realistic avatars from a single RGB image is an attractive yet challenging problem. To deal with challenging loose clothing or occlusion by interaction objects, we leverage powerful shape prior from 2D diffusion models pretrained on large datasets. Although 2D diffusion models demonstrate s…

Cited by 5SourcePDFScholar
2024

NRDF: Neural Riemannian Distance Fields for Learning Articulated Pose Priors

CVPR 2024highlight

Faithfully modeling the space of articulations is a crucial task that allows recovery and generation of realistic poses and remains a notorious challenge. To this end we introduce Neural Riemannian Distance Fields (NRDFs) data-driven priors modeling the space of plausible articulations represented a…

Cited by 11SourcePDFScholar
2024

Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering

CVPR 2024poster

We present Paint-it a text-driven high-fidelity texture map synthesis method for 3D meshes via neural re-parameterized texture optimization. Paint-it synthesizes texture maps from a text description by synthesis-through-optimization exploiting the Score-Distillation Sampling (SDS). We observe that d…

Cited by 44SourcePDFScholar
2024

Template Free Reconstruction of Human-object Interaction with Procedural Interaction Generation

CVPR 2024highlight

Reconstructing human-object interaction in 3D from a single RGB image is a challenging task and existing data driven methods do not generalize beyond the objects present in the carefully curated 3D interaction datasets. Capturing large-scale real data to learn strong interaction and 3D shape priors…

Cited by 14SourcePDFScholar
2023

NSF: Neural Surface Fields for Human Modeling from Monocular Depth

ICCV 2023poster

Obtaining personalized 3D animatable avatars from a monocular camera has several real world applications in gaming, virtual try-on, animation, and VR/XR, etc. However, it is very challenging to model dynamic and fine-grained clothing deformations from such sparse data. Existing methods for modeling…

Cited by 15PDFScholar
2023

Object Pop-Up: Can We Infer 3D Objects and Their Poses From Human Interactions Alone?

CVPR 2023poster

The intimate entanglement between objects affordances and human poses is of large interest, among others, for behavioural sciences, cognitive psychology, and Computer Vision communities. In recent years, the latter has developed several object-centric approaches: starting from items, learning pipeli…

2023

Visibility Aware Human-Object Interaction Tracking From Single RGB Camera

CVPR 2023poster

Capturing the interactions between humans and their environment in 3D is important for many applications in robotics, graphics, and vision. Recent works to reconstruct the 3D human and object from a single RGB image do not have consistent relative translation across frames because they assume a fixe…

Cited by 46SourcePDFScholar
2022

"CHORE: Contact, Human and Object REconstruction from a Single RGB Image"

ECCV 2022poster

"Most prior works in perceiving 3D humans from images reason human in isolation without their surroundings. However, humans are constantly interacting with the surrounding objects, thus calling for models that can reason about not only the human but also the object and their interaction. The problem…

2022

BEHAVE: Dataset and Method for Tracking Human Object Interactions

CVPR 2022poster

Modelling interactions between humans and objects in natural environments is central to many applications including gaming, virtual and mixed reality, as well as human behavior analysis and human-robot collaboration. This challenging operation scenario requires generalization to vast number of objec…

Cited by 213PDFcodeScholar
2022

Box2Mask: Weakly Supervised 3D Semantic Instance Segmentation Using Bounding Boxes

ECCV 2022poster

"Current 3D segmentation methods heavily rely on large-scale point-cloud datasets, which are notoriously laborious to annotate. Few attempts have been made to circumvent the need for dense per-point annotations. In this work, we look at weakly-supervised 3D semantic instance segmentation. The key id…

Cited by 76SourcePDFScholar
2022

COUCH: Towards Controllable Human-Chair Interactions

ECCV 2022poster

"Humans can interact with an object in the scene in many different ways, which are often associated with different modalities of contacting with the object. This creates a highly complex motion space that can be difficult to learn, particularly when synthesizing such human interactions in a controll…

Cited by 108SourcePDFScholar
2022

Learned Vertex Descent: A New Direction for 3D Human Model Fitting

ECCV 2022poster

"We propose a novel optimization-based paradigm for 3D human shape fitting on images. In contrast to existing approaches that directly regress the parameters of a low-dimensional statistical body model (e.g. SMPL) from input images, we propose training a deep network that, given solely image feature…

Cited by 40SourcePDFScholar
2022

Pose-NDF: Modeling Human Pose Manifolds with Neural Distance Fields

ECCV 2022poster

"We present Pose-NDF, a continuous model for plausible human poses based on neural distance fields (NDFs). Pose or motion priors are important for generating realistic new poses and for reconstructing accurate poses from noisy or partial observations. Pose-NDF learns a manifold of plausible poses as…

2022

Skeleton-Free Pose Transfer for Stylized 3D Characters

ECCV 2022poster

"We present the first method that automatically transfers poses between stylized 3D characters without skeletal rigging. In contrast to previous attempts to learn pose transformations on fixed or topology-equivalent skeleton templates, our method focuses on a novel scenario to handle skeleton-free c…

Cited by 41SourcePDFScholar
2022

TOCH: Spatio-Temporal Object-to-Hand Correspondence for Motion Refinement

ECCV 2022poster

"We present TOCH, a method for refining incorrect 3D hand-object interaction sequences using a data prior. Existing hand trackers, especially those that rely on very few cameras, often produce visually unrealistic results with hand-object intersection or missing contacts. Although correcting such er…

Cited by 56SourcePDFScholar
2021

D-NeRF: Neural Radiance Fields for Dynamic Scenes

CVPR 2021poster

Neural rendering techniques combining machine learning with geometric reasoning have arisen as one of the most promising approaches for synthesizing novel views of a scene from a sparse set of images. Among these, stands out the Neural radiance fields (NeRF), which trains a deep network to map 5D in…

Cited by 1585PDFScholar
2021

Human POSEitioning System (HPS): 3D Human Pose Estimation and Self-Localization in Large Scenes From Body-Mounted Sensors

CVPR 2021poster

We introduce (HPS) Human POSEitioning System, a method to recover the full 3D pose of a human registered with a 3D scan of the surrounding environment using wearable sensors. Using IMUs attached at the body limbs and a head mounted camera looking outwards, HPS fuses camera based self-localization wi…

Cited by 171PDFcodeScholar
2021

Neural-GIF: Neural Generalized Implicit Functions for Animating People in Clothing

ICCV 2021poster

We present Neural Generalized Implicit Functions(Neural-GIF), to animate people in clothing as a function of the body pose. Given a sequence of scans of a subject in various poses, we learn to animate the character for new poses. Existing methods have relied on template-based representations of the…

Cited by 128PDFcodeScholar
2021

SMPLicit: Topology-Aware Generative Model for Clothed People

CVPR 2021poster

In this paper we introduce SMPLicit, a novel generative model to jointly represent body pose, shape and clothing geometry. In contrast to existing learning-based approaches that require training specific models for each type of garment, SMPLicit can represent in a unified manner different garment to…

Cited by 219PDFcodeScholar
2021

Stereo Radiance Fields (SRF): Learning View Synthesis for Sparse Views of Novel Scenes

CVPR 2021poster

Recent neural view synthesis methods have achieved impressive quality and realism, surpassing classical pipelines which rely on multi-view reconstruction. State-of-the-Art methods, such as NeRF, are designed to learn a single scene with a neural network and require dense multi-view inputs. Testing o…

Cited by 261PDFScholar
2020

Combining Implicit Function Learning and Parametric Models for 3D Human Reconstruction

ECCV 2020poster

Implicit functions represented as deep learning approximations are powerful for reconstructing 3D surfaces. However, they can only produce static surfaces that are not controllable, which provides limited ability to modify the resulting model by editing its pose or shape parameters.Implicit function…

Cited by 236SourcePDFScholar
2020

DeepCap: Monocular Human Performance Capture Using Weak Supervision

CVPR 2020oral

Human performance capture is a highly important computer vision problem with many applications in movie production and virtual/augmented reality. Many previous performance capture approaches either required expensive multi-view setups or did not recover dense space-time coherent geometry with frame-…

Cited by 266PDFScholar
2020

Implicit Functions in Feature Space for 3D Shape Reconstruction and Completion

CVPR 2020poster

While many works focus on 3D reconstruction from images, in this paper, we focus on 3D shape reconstruction and completion from a variety of 3D inputs, which are deficient in some respect: low and high resolution voxels, sparse and dense point clouds, complete or incomplete. Processing of such 3D in…

Cited by 578PDFcodeScholar
2020

Learning to Dress 3D People in Generative Clothing

CVPR 2020poster

Three-dimensional human body models are widely used in the analysis of human pose and motion. Existing models, however, are learned from minimally-clothed 3D scans and thus do not generalize to the complexity of dressed people in common images and videos. Additionally, current models lack the expres…

Cited by 435PDFcodeScholar
2020

LoopReg: Self-supervised Learning of Implicit Surface Correspondences, Pose and Shape for 3D Human Mesh Registration

NeurIPS 2020oral

We address the problem of fitting 3D human models to 3D scans of dressed humans. Classical methods optimize both the data-to-model correspondences and the human model parameters (pose and shape), but are reliable only when initialised close to the solution. Some methods initialize the optimization b…

2020

NASA Neural Articulated Shape Approximation

ECCV 2020poster

Efficient representation of articulated objects such as human bodies is an important problem in computer vision and graphics. To efficiently simulate deformation, existing approaches represent 3D objects using polygonal meshes and deform them using skinning techniques. This paper introduces neural a…

Cited by 263SourcePDFScholar
2020

Neural Unsigned Distance Fields for Implicit Function Learning

NeurIPS 2020poster

In this work we target a learnable output representation that allows continuous, high resolution outputs of arbitrary shape. Recent works represent 3D surfaces implicitly with a Neural Network, thereby breaking previous barriers in resolution, and ability to represent diverse topologies. However, n…

Cited by 371SourcePDFScholar
2020

SIZER: A Dataset and Model for Parsing 3D Clothing and Learning Size Sensitive 3D Clothing

ECCV 2020poster

While models of 3D clothing learned from real data exist, no method can predict clothing deformation as a function of garment size. In this paper, we introduce SizerNet to predict 3D clothing conditioned on human body shape and garment size parameters, and ParserNet to infer garment meshes and shape…

2020

TailorNet: Predicting Clothing in 3D as a Function of Human Pose, Shape and Garment Style

CVPR 2020oral

In this paper, we present TailorNet, a neural model which predicts clothing deformation in 3D as a function of three factors: pose, shape and style (garment geometry), while retaining wrinkle detail. This goes beyond prior models, which are either specific to one style and shape, or generalize to di…

Cited by 344PDFcodeScholar
2019

AMASS: Archive of Motion Capture As Surface Shapes

ICCV 2019poster

Large datasets are the cornerstone of recent advances in computer vision using deep learning. In contrast, existing human motion capture (mocap) datasets are small and the motions limited, hampering progress on learning models of human motion. While there are many different datasets available, they…

Cited by 1579PDFcodeScholar
2019

In the Wild Human Pose Estimation Using Explicit 2D Features and Intermediate 3D Representations

CVPR 2019oral

Convolutional Neural Network based approaches for monocular 3D human pose estimation usually require a large amount of training images with 3D pose annotations. While it is feasible to provide 2D joint annotations for large corpora of in-the-wild images with humans, providing accurate 3D annotations…

Cited by 178PDFScholar
2019

Learning to Reconstruct People in Clothing From a Single RGB Camera

CVPR 2019poster

We present Octopus, a learning-based model to infer the personalized 3D shape of people from a few frames (1-8) of a monocular video in which the person is moving with a reconstruction accuracy of 4 to 5mm, while being orders of magnitude faster than previous methods. From semantic segmentation imag…

Cited by 380PDFcodeScholar
2019

Multi-Garment Net: Learning to Dress 3D People From Images

ICCV 2019poster

We present Multi-Garment Network (MGN), a method to predict body shape and clothing, layered on top of the SMPL model from a few frames (1-8) of a video. Several experiments demonstrate that this representation allows higher level of control when compared to single mesh or voxel representations of s…

Cited by 464PDFScholar
2019

SimulCap : Single-View Human Performance Capture With Cloth Simulation

CVPR 2019poster

This paper proposes a new method for live free-viewpoint human performance capture with dynamic details (e.g., cloth wrinkles) using a single RGBD camera. Our main contributions are: (i) a multi-layer representation of garments and body, and (ii) a physics-based performance capture procedure. We fir…

Cited by 125PDFScholar
2019

Tex2Shape: Detailed Full Human Body Geometry From a Single Image

ICCV 2019poster

We present a simple yet effective method to infer detailed full human body shape from only a single photograph. Our model can infer full-body shape including face, hair, and clothing including wrinkles at interactive frame-rates. Results feature details even on parts that are occluded in the input i…

Cited by 372PDFcodeScholar
2018

DoubleFusion: Real-Time Capture of Human Performances With Inner Body Shapes From a Single Depth Sensor

CVPR 2018poster

We propose DoubleFusion, a new real-time system that combines volumetric dynamic reconstruction with data-driven template fitting to simultaneously reconstruct detailed geometry, non-rigid motion and the inner human body shape from a single depth camera. One of the key contributions of this method i…

Cited by 371SourcePDFScholar
2018

Recovering Accurate 3D Human Pose in The Wild Using IMUs and a Moving Camera

ECCV 2018poster

In this work, we propose a method that combines a single hand-held camera and a set of Inertial Measurement Units (IMUs) attached at the body limbs to estimate accurate 3D poses in the wild. This poses many new challenges: the moving camera, heading drift, cluttered background, occlusions and many p…

Cited by 1257SourcePDFScholar
2018

Video Based Reconstruction of 3D People Models

CVPR 2018poster

This paper describes how to obtain accurate 3D body models and texture of arbitrary people from a single, monocular video in which a person is moving. Based on a parametric body model, we present a robust processing pipeline achieving 3D model fits with 5mm accuracy also for clothed people. Our main…

2017

Detailed, Accurate, Human Shape Estimation From Clothed 3D Scan Sequences

CVPR 2017spotlight

We address the problem of estimating human pose and body shape from 3D scans over time. Reliable estimation of 3D body shape is necessary for many applications including virtual try-on, health monitoring, and avatar creation for virtual reality. Scanning bodies in minimal clothing, however, presents…

Cited by 344PDFScholar