← Search

Michael J. Black

138 accepted papers

2026

MAMMA: Markerless Accurate Multi-person Motion Acquisition

CVPR 2026

We present MAMMA, a markerless motion-capture pipeline that accurately recovers SMPL-X parameters from multi-view video. Traditional motion-capture systems rely on physical markers. Although they offer high accuracy, their requirements of specialized hardware, manual marker placement, and extensive

Cited by 0SourcecodeScholar
2025

BEDLAM2.0: Synthetic humans and cameras in motion

NeurIPS 2025oral

Inferring 3D human motion from video remains a challenging problem with many applications. While traditional methods estimate the human in image coordinates, many applications require human motion to be estimated in world coordinates. This is particularly challenging when there is both human and cam…

Cited by 0SourceScholar
2025

Can Large Language Models Understand Symbolic Graphics Programs?

ICLR 2025spotlight

Against the backdrop of enthusiasm for large language models (LLMs), there is a growing need to scientifically assess their capabilities and shortcomings. This is nontrivial in part because it is difficult to find tasks which the models have not encountered during training. Utilizing symbolic graphi…

Cited by 11SourcePDFScholar
2025

ChatGarment: Garment Estimation, Generation and Editing via Large Language Models

CVPR 2025poster

We introduce ChatGarment, a novel approach that leverages large vision-language models (VLMs) to automate the estimation, generation, and editing of 3D garment sewing patterns from images or text descriptions. Unlike previous methods that often lack robustness and interactive editing capabilities, C…

Cited by 5SourcePDFScholar
2025

Contact-Aware Refinement of Human Pose Pseudo-Ground Truth via Bioimpedance Sensing

ICCV 2025poster

Capturing accurate 3D human pose in the wild would provide valuable data for training motion-generation and pose-estimation methods. While video-based estimation approaches have become increasingly accurate, they often fail in common scenarios involving self-contact, such as a hand touching the face…

Cited by 0SourcePDFScholar
2025

DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models

CVPR 2025poster

We address the task of generating 3D hair geometry from a single image, which is challenging due to the diversity of hairstyles and the lack of paired image-to-3D hair data. Previous methods are primarily trained on synthetic data and cope with the limited amount of such data by using low-dimensiona…

2025

ETCH: Generalizing Body Fitting to Clothed Humans via Equivariant Tightness

ICCV 2025poster

Fitting a body to a 3D clothed human point cloud is a common yet challenging task. Traditional optimization-based approaches use multi-stage pipelines that are sensitive to pose initialization, while recent learning-based methods often struggle with generalization across diverse poses and garment ty…

2025

Generative Zoo

ICCV 2025poster

The model-based estimation of 3D animal pose and shape from images enables computational modeling of animal behavior. Training models for this purpose requires large amounts of labeled image data with precise pose and shape annotations. However, capturing such data requires the use of multi-view or…

Cited by 0SourcePDFScholar
2025

HairFree: Compositional 2D Head Prior for Text-Driven 360° Bald Texture Synthesis

NeurIPS 2025poster

Synthesizing high-quality 3D head textures is crucial for gaming, virtual reality, and digital humans. Achieving seamless 360° textures typically requires expensive multi-view datasets with precise tracking. However, traditional methods struggle without back-view data or precise geometry, especially…

Cited by 0SourceScholar
2025

Im2Haircut: Single-view Strand-based Hair Reconstruction for Human Avatars

ICCV 2025poster

We present a novel approach for 3D hair reconstruction from single photographs based on a global hair prior combined with local optimization. Capturing strand-based hair geometry from single photographs is challenging due to the variety and geometric complexity of hairstyles and the lack of ground t…

Cited by 0SourcePDFScholar
2025

InterDyn: Controllable Interactive Dynamics with Video Diffusion Models

CVPR 2025poster

Predicting the dynamics of interacting objects is essential for both humans and intelligent systems. However, existing approaches are limited to simplified, toy settings and lack generalizability to complex, real-world environments. Recent advances in generative models have enabled the prediction of…

Cited by 2SourcePDFScholar
2025

InteractVLM: 3D Interaction Reasoning from 2D Foundational Models

CVPR 2025poster

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth ambiguities, and widely varying object shapes. Existing methods r…

2025

MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction

ICCV 2025poster

We present MoGA, a novel method to reconstruct high-fidelity 3D Gaussian avatars from a single-view image. The main challenge lies in inferring unseen appearance and geometric details while ensuring 3D consistency and realism. Most previous methods rely on 2D diffusion models to synthesize unseen vi…

2025

PICO: Reconstructing 3D People In Contact with Objects

CVPR 2025poster

Recovering 3D Human-Object Interaction (HOI) from single color images is challenging due to depth ambiguities, occlusions, and the huge variation in object shape and appearance. Thus, past work requires controlled settings such as known object shapes and contacts, and tackles only limited object cla…

Cited by 1SourcePDFScholar
2025

PRIMAL: Physically Reactive and Interactive Motor Model for Avatar Learning

ICCV 2025poster

We formulate the motor system of an interactive avatar as a generative motion model that can drive the body to move through 3D space in a perpetual, realistic, controllable, and responsive manner. Although human motion generation has been extensively studied, many existing methods lack the responsiv…

Cited by 0SourcePDFScholar
2025

PromptHMR: Promptable Human Mesh Recovery

CVPR 2025poster

Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxiliary "side information" that could enhance reconstruction accuracy in such challe…

Cited by 0SourcePDFScholar
2025

St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World

ICCV 2025poster

Dynamic 3D reconstruction and point tracking in videos are typically treated as separate tasks, despite their deep connection. We propose St4RTrack, a feed-forward frame- work that simultaneously reconstructs and tracks dynamic video content in a world coordinate frame from RGB in- puts. This is ach…

Cited by 0SourcePDFScholar
2024

ChatPose: Chatting about 3D Human Pose

CVPR 2024poster

We introduce ChatPose a framework employing Large Language Models (LLMs) to understand and reason about 3D human poses from images or textual descriptions. Our work is motivated by the human ability to intuitively understand postures from a single image or a brief description a process that intertwi…

2024

EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling

CVPR 2024poster

We propose EMAGE a framework to generate full-body human gestures from audio and masked gestures encompassing facial local body hands and global movements. To achieve this we first introduce BEAT2 (BEAT-SMPLX-FLAME) a new mesh-level holistic co-speech dataset. BEAT2 combines a MoShed SMPL-X body wit…

2024

Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

CVPR 2024poster

Existing methods for synthesizing 3D human gestures from speech have shown promising results but they do not explicitly model the impact of emotions on the generated gestures. Instead these methods directly output animations from speech without control over the expressed emotion. To address this lim…

2024

Explorative Inbetweening of Time and Space

ECCV 2024poster

"We introduce bounded generation as a generalized task to control video generation to synthesize arbitrary camera and subject motion based only on a given start and end frame. Our objective is to fully leverage the inherent generalization capability of an image-to-video model without additional trai…

Cited by 12SourcePDFScholar
2024

Generating Human Interaction Motions in Scenes with Text Control

ECCV 2024poster

"We present , a text-controlled scene-aware motion generation method based on denoising diffusion models. Previous text-to-motion methods focus on characters in isolation without considering scenes due to the limited availability of datasets that include motion, text descriptions, and interactive sc…

Cited by 40SourcePDFScholar
2024

Ghost on the Shell: An Expressive Representation of General 3D Shapes

ICLR 2024oral

The creation of photorealistic virtual worlds requires the accurate modeling of 3D surface geometry for a wide range of objects. For this, meshes are appealing since they enable 1) fast physics-based rendering with realistic material and lighting, 2) physical simulation, and 3) are memory-efficient…

Cited by 15SourcePDFScholar
2024

HIT: Estimating Internal Human Implicit Tissues from the Body Surface

CVPR 2024poster

The creation of personalized anatomical digital twins is important in the fields of medicine computer graphics sports science and biomechanics. To observe a subject's anatomy expensive medical devices (MRI or CT) are required and the creation of the digital model is often time-consuming and involves…

Cited by 3SourcePDFScholar
2024

HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video

CVPR 2024highlight

Since humans interact with diverse objects every day the holistic 3D capture of these interactions is important to understand and model human behaviour. However most existing methods for hand-object reconstruction from RGB either assume pre-scanned object templates or heavily rely on limited 3D hand…

2024

Human Hair Reconstruction with Strand-Aligned 3D Gaussians

ECCV 2024poster

"We introduce a new hair modeling method that uses a dual representation of classical hair strands and 3D Gaussians to produce accurate and realistic strand-based reconstructions from multi-view data. In contrast to recent approaches that leverage unstructured Gaussians to model human avatars, our m…

2024

Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization

ICLR 2024poster

Large foundation models are becoming ubiquitous, but training them from scratch is prohibitively expensive. Thus, efficiently adapting these powerful models to downstream tasks is increasingly important. In this paper, we study a principled finetuning paradigm -- Orthogonal Finetuning (OFT) -- for d…

Cited by 57SourcePDFScholar
2024

SCULPT: Shape-Conditioned Unpaired Learning of Pose-dependent Clothed and Textured Human Meshes

CVPR 2024poster

We present SCULPT a novel 3D generative model for clothed and textured 3D meshes of humans. Specifically we devise a deep neural network that learns to represent the geometry and appearance distribution of clothed human bodies. Training such a model is challenging as datasets of textured 3D meshes f…

Cited by 5SourcePDFScholar
2024

Synthesizing Environment-Specific People in Photographs

ECCV 2024poster

"We present ESP, a novel method for context-aware full-body generation, that enables photo-realistic synthesis and inpainting of people wearing clothing that is semantically appropriate for the scene depicted in an input photograph. ESP is conditioned on a 2D pose and contextual cues that are extrac…

Cited by 0SourcePDFScholar
2024

Text-Conditioned Generative Model of 3D Strand-based Human Hairstyles

CVPR 2024poster

We present HAAR a new strand-based generative model for 3D human hairstyles. Specifically based on textual inputs HAAR produces 3D hairstyles that could be used as production-level assets in modern computer graphics engines. Current AI-based generative models take advantage of powerful 2D priors to…

Cited by 3SourcePDFScholar
2024

TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

CVPR 2024poster

We address the problem of regressing 3D human pose and shape from a single image with a focus on 3D accuracy. The current best methods leverage large datasets of 3D pseudo-ground-truth (p-GT) and 2D keypoints leading to robust performance. With such methods however we observe a paradoxical decline i…

2024

VAREN: Very Accurate and Realistic Equine Network

CVPR 2024poster

Data-driven three-dimensional parametric shape models of the human body have gained enormous popularity both for the analysis of visual data and for the generation of synthetic humans. Following a similar approach for animals does not scale to the multitude of existing animal species not to mention…

Cited by 9SourcePDFScholar
2024

WANDR: Intention-guided Human Motion Generation

CVPR 2024poster

Synthesizing natural human motions that enable a 3D human avatar to walk and reach for arbitrary goals in 3D space remains an unsolved problem with many applications. Existing methods (data-driven or using reinforcement learning) are limited in terms of generalization and motion naturalness. A prima…

Cited by 12SourcePDFScholar
2024

WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion

CVPR 2024poster

The estimation of 3D human motion from video has progressed rapidly but current methods still have several key limitations. First most methods estimate the human in camera coordinates. Second prior work on estimating humans in global coordinates often assumes a flat ground plane and produces foot sl…

Cited by 75SourcePDFScholar
2023

3D Human Pose Estimation via Intuitive Physics

CVPR 2023poster

Estimating 3D humans from images often produces implausible bodies that lean, float, or penetrate the floor. Such methods ignore the fact that bodies are typically supported by the scene. A physics engine can be used to enforce physical plausibility, but these are not differentiable, rely on unreali…

Cited by 98SourcePDFScholar
2023

AG3D: Learning to Generate 3D Avatars from 2D Image Collections

ICCV 2023poster

While progress in 2D generative models of human appearance has been rapid, many applications require 3D avatars that can be animated and rendered. Unfortunately, most existing methods for learning generative models of 3D humans with diverse shape and appearance require 3D training data, which is lim…

Cited by 60PDFScholar
2023

ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

CVPR 2023poster

Humans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because there exist no datasets with ground-truth 3D annotations for…

2023

BEDLAM: A Synthetic Dataset of Bodies Exhibiting Detailed Lifelike Animated Motion

CVPR 2023highlight

We show, for the first time, that neural networks trained only on synthetic data achieve state-of-the-art accuracy on the problem of 3D human pose and shape (HPS) estimation from real images. Previous synthetic datasets have been small, unrealistic, or lacked realistic clothing. Achieving sufficient…

Cited by 158SourcePDFScholar
2023

BITE: Beyond Priors for Improved Three-D Dog Pose Estimation

CVPR 2023poster

We address the problem of inferring the 3D shape and pose of dogs from images. Given the lack of 3D training data, this problem is challenging, and the best methods lag behind those designed to estimate human shape and pose. To make progress, we attack the problem from multiple sides at once. First,…

Cited by 29SourcePDFScholar
2023

DECO: Dense Estimation of 3D Human-Scene Contact In The Wild

ICCV 2023oral

Understanding how humans use physical contact to interact with the world is key to enabling human-centric artificial intelligence. While inferring 3D contact is crucial for modeling realistic and physically-plausible human-object interactions, existing methods either focus on 2D, consider body joint…

Cited by 25PDFcodeScholar
2023

Detecting Human-Object Contact in Images

CVPR 2023poster

Humans constantly contact objects to move and perform tasks. Thus, detecting human-object contact is important for building human-centered artificial intelligence. However, there exists no robust method to detect contact between the body and the scene from an image, and there exists no dataset to le…

2023

ECON: Explicit Clothed Humans Optimized via Normal Integration

CVPR 2023highlight

The combination of deep learning, artist-curated scans, and Implicit Functions (IF), is enabling the creation of detailed, clothed, 3D humans from images. However, existing methods are far from perfect. IF-based methods recover free-form geometry, but produce disembodied limbs or degenerate shapes f…

2023

Generalizing Neural Human Fitting to Unseen Poses With Articulated SE(3) Equivariance

ICCV 2023oral

We address the problem of fitting a parametric human body model (SMPL) to point cloud data. Optimization based methods require careful initialization and are prone to becoming trapped in local optima. Learning-based methods address this but do not generalize well when the input pose is far from thos…

Cited by 14PDFScholar
2023

HOOD: Hierarchical Graphs for Generalized Modelling of Clothing Dynamics

CVPR 2023poster

We propose a method that leverages graph neural networks, multi-level message passing, and unsupervised training to enable real-time prediction of realistic clothing dynamics. Whereas existing methods based on linear blend skinning must be trained for specific garments, our method is agnostic to bod…

2023

Instant Multi-View Head Capture Through Learnable Registration

CVPR 2023poster

Existing methods for capturing datasets of 3D heads in dense semantic correspondence are slow and commonly address the problem in two separate steps; multi-view stereo (MVS) reconstruction followed by non-rigid registration. To simplify this process, we introduce TEMPEH (Towards Estimation of 3D Mes…

2023

MIME: Human-Aware 3D Scene Generation

CVPR 2023poster

Generating realistic 3D worlds occupied by moving humans has many applications in games, architecture, and synthetic data creation. But generating such scenes is expensive and labor intensive. Recent work generates human poses and motions given a 3D scene. Here, we take the opposite approach and gen…

2023

MeshDiffusion: Score-based Generative 3D Mesh Modeling

ICLR 2023top-25%

We consider the task of generating realistic 3D shapes, which is useful for a variety of applications such as automatic scene generation and physical simulation. Compared to other 3D representations like voxels and point clouds, meshes are more desirable in practice, because (1) they enable easy and…

2023

Pairwise Similarity Learning is SimPLE

ICCV 2023poster

In this paper, we focus on a general yet important learning problem, pairwise similarity learning (PSL). PSL subsumes a wide range of important applications, such as open-set face recognition, speaker verification, image retrieval and person re-identification. The goal of PSL is to learn a pairwise…

Cited by 10PDFcodeScholar
2023

PointAvatar: Deformable Point-Based Head Avatars From Videos

CVPR 2023poster

The ability to create realistic animatable and relightable head avatars from casual video sequences would open up wide ranging applications in communication and entertainment. Current methods either build on explicit 3D morphable meshes (3DMM) or exploit neural implicit representations. The former a…

2023

Reconstructing Signing Avatars From Video Using Linguistic Priors

CVPR 2023poster

Sign language (SL) is the primary method of communication for the 70 million Deaf people around the world. Video dictionaries of isolated signs are a core SL learning tool. Replacing these with 3D avatars can aid learning and enable AR/VR applications, improving access to technology and online media…

Cited by 15SourcePDFScholar
2023

SINC: Spatial Composition of 3D Human Motions for Simultaneous Action Generation

ICCV 2023poster

Our goal is to synthesize 3D human motions given textual inputs describing simultaneous actions, for example `waving hand' while `walking' at the same time. We refer to generating such simultaneous movements as performing `spatial compositions'. In contrast to `temporal compositions' that seek to tr…

Cited by 48PDFScholar
2023

SmartMocap: Joint Estimation of Human and Camera Motion Using Uncalibrated RGB Cameras

RA-L 2023

Markerless human motion capture (mocap) from multiple RGB cameras is a widely studied problem. Existing methods either need calibrated cameras or calibrate them relative to a static camera, which acts as the reference frame for the mocap system. The calibration step has to be done a priori for every

Cited by 13SourcecodeScholar
2023

TRACE: 5D Temporal Regression of Avatars With Dynamic Cameras in 3D Environments

CVPR 2023poster

Although the estimation of 3D human pose and shape (HPS) is rapidly progressing, current methods still cannot reliably estimate moving humans in global coordinates, which is critical for many applications. This is particularly challenging when the camera is also moving, entangling human and camera m…

2023

Viewpoint-Driven Formation Control of Airships for Cooperative Target Tracking

RA-L 2023

For tracking and motion capture (MoCap) of animals in their natural habitat, a formation of safe and silent aerial platforms, such as airships with on-board cameras, is well suited. In our prior work we derived formation properties for optimal MoCap, which include maintaining constant angular separa

Cited by 11SourcecodeScholar
2022

Accurate 3D Body Shape Regression Using Metric and Semantic Attributes

CVPR 2022oral

While methods that regress 3D human meshes from images have progressed rapidly, the estimated body shapes often do not capture the true human shape. This is problematic since, for many applications, accurate body shape is as important as pose. The key reason that body shape accuracy lags pose accura…

Cited by 72PDFcodeScholar
2022

AirPose: Multi-View Fusion Network for Aerial 3D Human Pose and Shape Estimation

RA-L 2022

In this letter, we present a novel markerless 3D human motion capture (MoCap) system for unstructured, outdoor environments that uses a team of autonomous unmanned aerial vehicles (UAVs) with on-board RGB cameras and computation. Existing methods are limited by calibrated cameras and off-line proces

Cited by 32SourcecodeScholar
2022

BARC: Learning To Regress 3D Dog Shape From Images by Exploiting Breed Information

CVPR 2022poster

Our goal is to recover the 3D shape and pose of dogs from a single image. This is a challenging task because dogs exhibit a wide range of shapes and appearances, and are highly articulated. Recent work has proposed to directly regress the SMAL animal model, with additional limb scale parameters, fro…

Cited by 52PDFScholar
2022

Capturing and Inferring Dense Full-Body Human-Scene Contact

CVPR 2022poster

Inferring human-scene contact (HSC) is the first step toward understanding how humans interact with their surroundings. While detecting 2D human-object interaction (HOI) and reconstructing 3D human pose and shape (HPS) have enjoyed significant progress, reasoning about 3D human-scene contact from a…

Cited by 150PDFScholar
2022

Deep Residual Reinforcement Learning based Autonomous Blimp Control

IROS 2022poster

Blimps are well suited to perform long-duration aerial tasks as they are energy efficient, relatively silent and safe. To address the blimp navigation and control task, in previous work we developed a hardware and software-in-the-loop framework and a PID-based controller for large blimps in the pres…

Cited by 14SourcecodeScholar
2022

GOAL: Generating 4D Whole-Body Motion for Hand-Object Grasping

CVPR 2022poster

Generating digital humans that move realistically has many applications and is widely studied, but existing methods focus on the major limbs of the body, ignoring the hands and head. Hands have been separately studied, but the focus has been on generating realistic static grasps of objects. To synth…

Cited by 124PDFcodeScholar
2022

Human-Aware Object Placement for Visual Environment Reconstruction

CVPR 2022poster

Humans are in constant contact with the world as they move through it and interact with it. This contact is a vital source of information for understanding 3D humans, 3D scenes, and the interactions between them. In fact, we demonstrate that these human-scene interactions (HSIs) can be leveraged to…

Cited by 70PDFcodeScholar
2022

I M Avatar: Implicit Morphable Head Avatars From Videos

CVPR 2022oral

Traditional 3D morphable face models (3DMMs) provide fine-grained control over expression but cannot easily capture geometric and appearance details. Neural volumetric representations approach photorealism but are hard to animate and do not generalize well to unseen expressions. To tackle this probl…

Cited by 254PDFcodeScholar
2022

ICON: Implicit Clothed Humans Obtained From Normals

CVPR 2022poster

Current methods for learning realistic and animatable 3D clothed avatars need either posed 3D scans or 2D images with carefully controlled user poses. In contrast, our goal is to learn the avatar from only 2D images of people in unconstrained poses. Given a set of images, our method estimates a deta…

Cited by 343PDFcodeScholar
2022

Putting People in Their Place: Monocular Regression of 3D People in Depth

CVPR 2022poster

Given an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene co…

Cited by 180PDFcodeScholar
2022

SUPR: A Sparse Unified Part-Based Human Representation

ECCV 2022poster

"Statistical 3D shape models of the head, hands, and full body are widely used in computer vision and graphics. Despite their wide use, we show that existing models of the head and hands fail to capture the full range of motion for these parts. Moreover, existing work largely ignores the feet, which…

2022

TEMOS: Generating Diverse Human Motions from Textual Descriptions

ECCV 2022poster

"We address the problem of generating diverse 3D human motions from textual descriptions. This challenging task requires joint modeling of both modalities: understanding and extracting useful human-centric information from the text, and then generating plausible and realistic sequences of human pose…

2022

Towards Racially Unbiased Skin Tone Estimation via Scene Disambiguation

ECCV 2022poster

"Virtual facial avatars will play an increasingly important role in immersive communication, games and the metaverse, and it is therefore critical that they be inclusive. This requires accurate recovery of the albedo, regardless of age, sex, or ethnicity. While significant progress has been made on…

Cited by 35SourcePDFScholar
2022

gDNA: Towards Generative Detailed Neural Avatars

CVPR 2022poster

To make 3D human avatars widely available, we must be able to generate a variety of 3D virtual humans with varied identities and shapes in arbitrary poses. This task is challenging due to the diversity of clothed body shapes, their complex articulations, and the resulting rich, yet stochastic geomet…

Cited by 84PDFScholar
2021

AGORA: Avatars in Geography Optimized for Regression Analysis

CVPR 2021poster

While the accuracy of 3D human pose estimation from images has steadily improved on benchmark datasets, the best methods still fail in many real-world scenarios. This suggests that there is a domain gap between current datasets and common scenes containing people. To obtain ground-truth 3D pose, cur…

Cited by 245PDFcodeScholar
2021

BABEL: Bodies, Action and Behavior With English Labels

CVPR 2021poster

Understanding the semantics of human movement -- the what, how and why of the movement -- is an important problem that requires datasets of human actions with semantic labels. Existing datasets take one of two approaches. Large-scale video datasets contain many action labels but do not contain groun…

Cited by 236PDFcodeScholar
2021

Learning Realistic Human Reposing Using Cyclic Self-Supervision With 3D Shape, Pose, and Appearance Consistency

ICCV 2021poster

Synthesizing images of a person in novel poses from a single image is a highly ambiguous task. Most existing approaches require paired training images; i.e. images of the same person with the same clothing in different poses. However, obtaining sufficiently large datasets with paired data is challen…

Cited by 20PDFScholar
2021

Learning To Regress Bodies From Images Using Differentiable Semantic Rendering

ICCV 2021poster

Learning to regress 3D human body shape and pose (e.g. SMPL parameters) from monocular images typically exploits losses on 2D keypoints, silhouettes, and/or part-segmentation when 3D training data is not available. Such losses, however, are limited because 2D keypoints do not supervise body shape an…

Cited by 65PDFcodeScholar
2021

Monocular, One-Stage, Regression of Multiple 3D People

ICCV 2021poster

This paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a On…

Cited by 327PDFcodeScholar
2021

PARE: Part Attention Regressor for 3D Human Body Estimation

ICCV 2021poster

Despite significant progress, we show that state of the art 3D human pose and shape estimation methods remain sensitive to partial occlusion and can produce dramatically wrong predictions although much of the body is observable. To address this, we introduce a soft attention mechanism, called the Pa…

Cited by 480PDFcodeScholar
2021

Populating 3D Scenes by Learning Human-Scene Interaction

CVPR 2021poster

Humans live within a 3D space and constantly interact with it to perform tasks. Such interactions involve physical contact between surfaces that is semantically meaningful. Our goal is to learn how humans interact with scenes and leverage this to enable virtual characters to do the same. To that end…

Cited by 167PDFcodeScholar
2021

SCALE: Modeling Clothed Humans with a Surface Codec of Articulated Local Elements

CVPR 2021poster

Learning to model and reconstruct humans in clothing is challenging due to articulation, non-rigid deformation, and varying clothing types and topologies. To enable learning, the choice of representation is the key. Recent work uses neural networks to parameterize local surface elements. This approa…

Cited by 114PDFcodeScholar
2021

SCANimate: Weakly Supervised Learning of Skinned Clothed Avatar Networks

CVPR 2021poster

We present SCANimate, an end-to-end trainable framework that takes raw 3D scans of a clothed human and turns them into an animatable avatar. These avatars are driven by pose parameters and have realistic clothing that moves and deforms naturally. SCANimate does not rely on a customized mesh template…

Cited by 264PDFcodeScholar
2021

SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes

ICCV 2021poster

Neural implicit surface representations have emerged as a promising paradigm to capture 3D shapes in a continuous and resolution-independent manner. However, adapting them to articulated shapes is non-trivial. Existing approaches learn a backward warp field that maps deformed to canonical points. Ho…

Cited by 258PDFcodeScholar
2021

SPEC: Seeing People in the Wild With an Estimated Camera

ICCV 2021poster

Due to the lack of camera parameter information for in-the-wild images, existing 3D human pose and shape (HPS) estimation methods make several simplifying assumptions: weak-perspective projection, large constant focal length, and zero camera rotation. These assumptions often do not hold and we show,…

Cited by 160PDFcodeScholar
2020

AirCapRL: Autonomous Aerial Human Motion Capture Using Deep Reinforcement Learning

RA-L 2020

In this letter, we introduce a deep reinforcement learning (RL) based multi-robot formation controller for the task of autonomous aerial human motion capture (MoCap). We focus on vision-based MoCap, where the objective is to estimate the trajectory of body pose and shape of a single moving person us

Cited by 33SourceScholar
2020

GRAB: A Dataset of Whole-Body Human Grasping of Objects

ECCV 2020poster

Training computers to understand, model, and synthesize human grasping requires a rich dataset containing complex 3D object shapes, detailed contact information, hand pose and shape, and the 3D body motion over time. While ""grasping"" is commonly thought of as a single hand stably lifting an object…

2020

Learning to Dress 3D People in Generative Clothing

CVPR 2020poster

Three-dimensional human body models are widely used in the analysis of human pose and motion. Existing models, however, are learned from minimally-clothed 3D scans and thus do not generalize to the complexity of dressed people in common images and videos. Additionally, current models lack the expres…

Cited by 435PDFcodeScholar
2020

Monocular Expressive Body Regression through Body-Driven Attention

ECCV 2020poster

To understand how people look, interact, or perform tasks, we need to quickly and accurately capture their 3D body, face, and hands together from an RGB image. Most existing methods focus only on parts of the body. A few recent approaches reconstruct full expressive 3D humans from images using 3D bo…

2020

STAR: Sparse Trained Articulated Human Body Regressor

ECCV 2020poster

The SMPL body model is widely used for the estimation, synthesis, and analysis of 3D human pose and shape. While popular, we show that SMPL has several limitations and introduce STAR, which is quantitatively and qualitatively superior to SMPL. First, SMPL has a huge number of parameters resulting fr…

2020

VIBE: Video Inference for Human Body Pose and Shape Estimation

CVPR 2020poster

Human motion is fundamental to understanding behavior. Despite progress on single-image 3D pose and shape estimation, existing video-based state-of-the-art methods fail to produce accurate and natural motion sequences due to a lack of ground-truth 3D motion data for training. To address this problem…

Cited by 1216PDFcodeScholar
2019

AMASS: Archive of Motion Capture As Surface Shapes

ICCV 2019poster

Large datasets are the cornerstone of recent advances in computer vision using deep learning. In contrast, existing human motion capture (mocap) datasets are small and the motions limited, hampering progress on learning models of human motion. While there are many different datasets available, they…

Cited by 1579PDFcodeScholar
2019

Active Perception Based Formation Control for Multiple Aerial Vehicles

RA-L 2019

We present a novel robotic front-end for autonomous aerial motion-capture (mocap) in outdoor environments. In previous work, we presented an approach for cooperative detection and tracking (CDT) of a subject using multiple micro-aerial vehicles (MAVs). However, it did not ensure optimal view-point c

Cited by 69SourceScholar
2019

Capture, Learning, and Synthesis of 3D Speaking Styles

CVPR 2019poster

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we introduce a unique 4D face dataset with about 29 minutes of 4D…

Cited by 426PDFcodeScholar
2019

Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation

CVPR 2019poster

We address the unsupervised learning of several interconnected problems in low-level vision: single view depth prediction, camera motion estimation, optical flow, and segmentation of a video into the static scene and moving regions. Our key insight is that these four fundamental vision problems are…

Cited by 742PDFcodeScholar
2019

Expressive Body Capture: 3D Hands, Face, and Body From a Single Image

CVPR 2019oral

To facilitate the analysis of human actions, interactions and emotions, we compute a 3D model of human body pose, hand pose, and facial expression from a single monocular image. To achieve this, we use thousands of 3D scans to train a new, unified, 3D model of the human body, SMPL-X, that extends SM…

Cited by 2080PDFcodeScholar
2019

Learning Joint Reconstruction of Hands and Manipulated Objects

CVPR 2019poster

Estimating hand-object manipulations is essential for in- terpreting and imitating human actions. Previous work has made significant progress towards reconstruction of hand poses and object shapes in isolation. Yet, reconstructing hands and objects during manipulation is a more challeng- ing task du…

Cited by 631PDFScholar
2019

Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the Loop

ICCV 2019poster

Model-based human pose estimation is currently approached through two different paradigms. Optimization-based methods fit a parametric body model to 2D observations in an iterative manner, leading to accurate image-model alignments, but are often slow and sensitive to the initialization. In contrast…

Cited by 1232PDFScholar
2019

Learning to Regress 3D Face Shape and Expression From an Image Without 3D Supervision

CVPR 2019poster

The estimation of 3D face shape from a single image must be robust to variations in lighting, head pose, expression, facial hair, makeup, and occlusions. Robustness requires a large training set of in-the-wild images, which by construction, lack ground truth 3D shape. To train a network without any…

Cited by 363PDFScholar
2019

Markerless Outdoor Human Motion Capture Using Multiple Autonomous Micro Aerial Vehicles

ICCV 2019poster

Capturing human motion in natural scenarios means moving motion capture out of the lab and into the wild. Typical approaches rely on fixed, calibrated, cameras and reflective markers on the body, significantly limiting the motions that can be captured. To make motion capture truly unconstrained, we…

Cited by 43PDFScholar
2019

Resolving 3D Human Pose Ambiguities With 3D Scene Constraints

ICCV 2019poster

To understand and analyze human behavior, we need to capture humans moving in, and interacting with, the world. Most existing methods perform 3D human pose estimation without explicitly considering the scene. We observe however that the world constrains the body and vice-versa. To motivate this, we…

Cited by 359PDFcodeScholar
2019

Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"

ICCV 2019poster

We present the first method to perform automatic 3D pose, shape and texture capture of animals from images acquired in-the-wild. In particular, we focus on the problem of capturing 3D information about Grevy's zebras from a collection of images. The Grevy's zebra is one of the most endangered specie…

Cited by 187PDFcodeScholar
2018

Deep Neural Network-Based Cooperative Visual Tracking Through Multiple Micro Aerial Vehicles

RA-L 2018

Multicamera tracking of humans and animals in outdoor environments is a relevant and challenging problem. Our approach to it involves a team of cooperating microaerial vehicles (MAVs) with on-board cameras only. Deep neural networks (DNNs) often fail at detecting small-scale objects or those that ar

Cited by 64SourceScholar
2018

End-to-End Recovery of Human Shape and Pose

CVPR 2018poster

We describe Human Mesh Recovery (HMR), an end-to-end framework for reconstructing a full 3D mesh of a human body from a single RGB image. In contrast to most current methods that compute 2D or 3D joint locations, we produce a richer and more useful mesh representation that is parameterized by shape…

2018

Generating 3D Faces using Convolutional Mesh Autoencoders

ECCV 2018poster

Learned 3D representations of human faces are useful for computer vision problems such as 3D face tracking and reconstruction from images, as well as graphics applications such as character generation and animation. Traditional models learn a latent representation of a face using linear subspaces or…

2018

Lions and Tigers and Bears: Capturing Non-Rigid, 3D, Articulated Shape From Images

CVPR 2018poster

Animals are widespread in nature and the analysis of their shape and motion is important in many fields and industries. Modeling 3D animal shape, however, is difficult because the 3D scanning methods used to capture human shape are not applicable to wild animals or natural settings. Consequently, we…

Cited by 154SourcePDFScholar
2018

Recovering Accurate 3D Human Pose in The Wild Using IMUs and a Moving Camera

ECCV 2018poster

In this work, we propose a method that combines a single hand-held camera and a set of Inertial Measurement Units (IMUs) attached at the body limbs to estimate accurate 3D poses in the wild. This poses many new challenges: the moving camera, heading drift, cluttered background, occlusions and many p…

Cited by 1257SourcePDFScholar
2017

3D Menagerie: Modeling the 3D Shape and Pose of Animals

CVPR 2017spotlight

There has been significant work on learning realistic, articulated, 3D models of the human body. In contrast, there are few such models of animals, despite many applications. The main challenge is that animals are much less cooperative than humans. The best human body models are learned from thousan…

Cited by 491PDFScholar
2017

Deep Representation Learning for Human Motion Prediction and Classification

CVPR 2017poster

Generative models of 3D human motion are often restricted to a small number of activities and can therefore not generalize well to novel movements or applications. In this work we propose a deep learning framework for human motion capture data that learns a generic representation from a large co…

Cited by 519PDFScholar
2017

Detailed, Accurate, Human Shape Estimation From Clothed 3D Scan Sequences

CVPR 2017spotlight

We address the problem of estimating human pose and body shape from 3D scans over time. Reliable estimation of 3D body shape is necessary for many applications including virtual try-on, health monitoring, and avatar creation for virtual reality. Scanning bodies in minimal clothing, however, presents…

Cited by 344PDFScholar
2017

Learning From Synthetic Humans

CVPR 2017poster

Estimating human pose, shape, and motion from images and video are fundamental challenges with many applications. Recent advances in 2D human pose estimation use large amounts of manually-labeled training data for learning convolutional neural networks (CNNs). Such data is time consuming to acquire…

Cited by 1234PDFScholar
2017

Slow Flow: Exploiting High-Speed Cameras for Accurate and Diverse Optical Flow Reference Data

CVPR 2017oral

Existing optical flow datasets are limited in size and variability due to the difficulty of capturing dense ground truth. In this paper, we tackle this problem by tracking pixels through densely sampled space-time volumes recorded with a high-speed video camera. Our model exploits the linearity of s…

Cited by 97PDFScholar
2017

Unite the People: Closing the Loop Between 3D and 2D Human Representations

CVPR 2017poster

3D models provide a common ground for different representations of human bodies. In turn, robust 2D estimation has proven to be a powerful tool to obtain 3D fits "in-the-wild". However, depending on the level of detail, it can be hard to impossible to acquire labeled data for training 2D estimators…

Cited by 680PDFScholar
2016

Optical Flow With Semantic Segmentation and Localized Layers

CVPR 2016spotlight

Existing optical flow methods make generic, spatially homogeneous, assumptions about the spatial structure of the flow. In reality, optical flow varies across an image depending on object class. Simply put, different objects move differently. Here we exploit recent advances in static semantic scene…

Cited by 251PDFScholar
2016

Patches, Planes and Probabilities: A Non-Local Prior for Volumetric 3D Reconstruction

CVPR 2016poster

In this paper, we propose a non-local structured prior for volumetric multi-view 3D reconstruction. Towards this goal, we present a novel Markov random field model based on ray potentials in which assumptions about large 3D surface patches such as planarity or Manhattan world constraints can be effi…

Cited by 42PDFScholar
2015

Detailed Full-Body Reconstructions of Moving People From Monocular RGB-D Sequences

ICCV 2015poster

We accurately estimate the 3D geometry and appearance of the human body from a monocular RGB-D sequence of a user moving freely in front of the sensor. Range data in each frame is first brought into alignment with a multi-resolution 3D body model in a coarse-to-fine process. The method then uses geo…

Cited by 259PDFcodeScholar
2015

Efficient Sparse-to-Dense Optical Flow Estimation Using a Learned Basis and Layers

CVPR 2015poster

We address the elusive goal of estimating optical flow both accurately and efficiently by adopting a sparse-to-dense approach. Given a set of sparse matches, we regress to dense optical flow using a learned set of full-frame basis flow fields. We learn the principal components of natural flow fields…

Cited by 224SourcePDFScholar