← Search

Georgios Pavlakos

44 accepted papers

2026

HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image

CVPR 2026

In this paper, we present HumanNOVA, a photorealistic, universal, and rapid model for generating 3D human avatars from a single RGB image. Achieving both photorealism and generalization is challenging due to the scarcity of diverse, high-quality 3D human data. To address this, we build a scalable da

Cited by 0SourcecodeScholar
2026

Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos

CVPR 2026

Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate. We observe that this limitation stems from the absence of reliable high-order temporal cues--velocity and acceleration--which are essential for re

Cited by 0SourceScholar
2026

Recovering Physically Plausible Human-Object Interactions from Monocular Videos

CVPR 2026

In this paper, we present a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kinematic-based approaches produce visually plausible motion, they often result in physical artifacts such as interpenetration and object floating. To overcome

Cited by 0SourceScholar
2026

Searching in Space and Time: Unified Memory-Action Loops for Open-World Object Retrieval

ICRA 2026poster

Service robots must retrieve objects in dynamic, open-world settings where requests may reference attributes (“the red mug”), spatial context (“the mug on the table”), or past states (“the mug that was here yesterday”). Existing approaches capture only parts of this problem: scene graphs capture spa…

2025

Atlas Gaussians Diffusion for 3D Generation

ICLR 2025spotlight

Using the latent diffusion model has proven effective in developing novel 3D generation techniques. To harness the latent diffusion model, a key challenge is designing a high-fidelity and efficient representation that links the latent space and the 3D space. In this paper, we introduce Atlas Gaussia…

2025

COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning

CoRL 2025poster

In this work, we study the problem of data retrieval for few-shot imitation learning: select data from a large dataset to train a performant policy for a specific task, given only a few target demonstrations. Prior methods retrieve data using a single-feature distance heuristic, assuming that the be…

Cited by 0SourcecodeScholar
2025

Estimating Body and Hand Motion in an Ego-sensed World

CVPR 2025highlight

We present EgoAllo, a system for human motion estimation from a head-mounted device. Using only egocentric SLAM poses and images, EgoAllo guides sampling from a conditional diffusion model to estimate 3D body pose, height, and hand parameters that capture a device wearer's actions in the allocentric…

Cited by 5SourcePDFScholar
2025

ExpertAF: Expert Actionable Feedback from Video

CVPR 2025poster

Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demonstrations, leaving the burden of knowing what to do differently on the user. We introduce a novel method to generate act…

Cited by 3SourcePDFScholar
2025

MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

CVPR 2025poster

We propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes - over 50 times larger than the prior real dataset DL3DV - dramatically scaling the training data. To enable scalable data…

Cited by 1SourcePDFScholar
2025

RayZer: A Self-supervised Large View Synthesis Model

ICCV 2025poster

We present RayZer, a self-supervised multi-view 3D Vision model trained without any 3D supervision, i.e., camera poses and scene geometry, while exhibiting emerging 3D awareness. Concretely, RayZer takes unposed and uncalibrated images as input, recovers camera parameters, reconstructs a scene repre…

Cited by 0SourcePDFScholar
2025

Real3D: Towards Scaling Large Reconstruction Models with Real Images

ICCV 2025poster

Training single-view Large Reconstruction Models (LRMs) follows the fully supervised route, requiring multi-view supervision. However, the multi-view data typically comes from synthetic 3D assets, which are hard to scale further and are not representative of the distribution of real-world object sha…

Cited by 0SourcePDFScholar
2025

Reconstructing Humans with a Biomechanically Accurate Skeleton

CVPR 2025poster

In this paper, we introduce a method for reconstructing 3D humans from a single image using a biomechanically accurate skeleton model. To achieve this, we train a transformer that takes an image as input and estimates the parameters of the model. Due to the lack of training data for this task, we bu…

2024

CoFie: Learning Compact Neural Surface Representations with Coordinate Fields

NeurIPS 2024poster

This paper introduces CoFie, a novel local geometry-aware neural surface representation. CoFie is motivated by the theoretical analysis of local SDFs with quadratic approximation. We find that local shapes are highly compressive in an aligned coordinate frame defined by the normal and tangent direct…

2024

Expressive Gaussian Human Avatars from Monocular RGB Video

NeurIPS 2024poster

Nuanced expressiveness, especially through detailed hand and facial expressions, is pivotal for enhancing the realism and vitality of digital human representations. In this work, we aim to learn expressive human avatars from a monocular RGB video; a setting that introduces new challenges in capturin…

2024

GART: Gaussian Articulated Template Models

CVPR 2024highlight

We introduce Gaussian Articulated Template Model (GART) an explicit efficient and expressive representation for non-rigid articulated subject capturing and rendering from monocular videos. GART utilizes a mixture of moving 3D Gaussians to explicitly approximate a deformable subject's geometry and ap…

Cited by 94SourcePDFScholar
2024

Generative Proxemics: A Prior for 3D Social Interaction from Images

CVPR 2024poster

Social interaction is a fundamental aspect of human behavior and communication. The way individuals position themselves in relation to others also known as proxemics conveys social cues and affects the dynamics of social interaction. Reconstructing such interaction from images presents challenges be…

2024

MultiPhys: Multi-Person Physics-aware 3D Motion Estimation

CVPR 2024poster

We introduce MultiPhys a method designed for recovering multi-person motion from monocular videos. Our focus lies in capturing coherent spatial placement between pairs of individuals across varying degrees of engagement. MultiPhys being physically aware exhibits robustness to jittering and occlusion…

Cited by 5SourcePDFScholar
2024

OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation

CoRL 2024poster

We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for execution. At the heart of our approach is object-aware retargeting,…

Cited by 33SourcecodeScholar
2024

Reconstructing Hands in 3D with Transformers

CVPR 2024poster

We present an approach that can reconstruct hands in 3D from monocular input. Our approach for Hand Mesh Recovery HaMeR follows a fully transformer-based architecture and can analyze hands with significantly increased accuracy and robustness compared to previous work. The key to HaMeR's success lies…

2023

Decoupling Human and Camera Motion From Videos in the Wild

CVPR 2023poster

We propose a method to reconstruct global human trajectories from videos in the wild. Our optimization method decouples the camera and human motion, which allows us to place people in the same world coordinate frame. Most existing methods do not model the camera motion; methods that rely on the back…

2023

Humans in 4D: Reconstructing and Tracking Humans with Transformers

ICCV 2023poster

We present an approach to reconstruct humans and track them over time. At the core of our approach, we propose a fully "transformerized" version of a network for human mesh recovery. This network, HMR 2.0, advances the state of the art and shows the capability to analyze unusual poses that have in t…

Cited by 221PDFcodeScholar
2023

Learning Articulated Shape With Keypoint Pseudo-Labels From Web Images

CVPR 2023poster

This paper shows that it is possible to learn models for monocular 3D reconstruction of articulated objects (e.g. horses, cows, sheep), using as few as 50-150 images labeled with 2D keypoints. Our proposed approach involves training category-specific keypoint estimators, generating 2D keypoint pseud…

Cited by 8SourcePDFScholar
2023

On the Benefits of 3D Pose and Tracking for Human Action Recognition

CVPR 2023poster

In this work we study the benefits of using tracking and 3D poses for action recognition. To achieve this, we take the Lagrangian view on analysing actions over a trajectory of human motion rather than at a fixed point in space. Taking this stand allows us to use the tracklets of people to predict t…

2022

Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language Recognition

ICASSP 2022accepted

We address the challenging problem of continuous sign language recognition (CSLR) from RGB videos, proposing a novel deep-learning framework that employs spatio-temporal graph convolutional networks (ST-GCNs), which operate on multiple, appropriately fused feature streams, capturing the signer’s pos…

Cited by 0SourceScholar
2022

The One Where They Reconstructed 3D Humans and Environments in TV Shows

ECCV 2022poster

"TV shows depict a wide variety of human behaviors and have been studied extensively for their potential to be a rich source of data for many applications. However, the majority of the existing work focuses on 2D recognition tasks. In this paper, we make the observation that there is a certain persi…

Cited by 28SourcePDFScholar
2022

Tracking People by Predicting 3D Appearance, Location and Pose

CVPR 2022oral

We present an approach for tracking people in monocular videos by predicting their future 3D representations. To achieve this, we first lift people to 3D from a single frame in a robust manner. This lifting includes information about the 3D pose of the person, their location in the 3D space, and the…

Cited by 74PDFcodeScholar
2021

Independent Sign Language Recognition with 3d Body, Hands, and Face Reconstruction

ICASSP 2021accepted

Independent Sign Language Recognition is a complex visual recognition problem that combines several challenging tasks of Computer Vision due to the necessity to exploit and fuse information from hand gestures, body features and facial expressions. While many state-of-the-art works have managed to de…

Cited by 0SourceScholar
2021

Probabilistic Modeling for Human Mesh Recovery

ICCV 2021poster

This paper focuses on the problem of 3D human reconstruction from 2D evidence. Although this is an inherently ambiguous problem, the majority of recent works avoid the uncertainty modeling and typically regress a single estimate for a given input. In contrast to that, in this work, we propose to emb…

Cited by 212PDFcodeScholar
2021

Tracking People with 3D Representations

NeurIPS 2021poster

We present a novel approach for tracking multiple people in video. Unlike past approaches which employ 2D representations, we focus on using 3D representations of people, located in three-dimensional space. To this end, we develop a method, Human Mesh and Appearance Recovery (HMAR) which in addition…

2020

Coherent Reconstruction of Multiple Humans From a Single Image

CVPR 2020poster

In this work, we address the problem of multi-person 3D pose estimation from a single image. A typical regression approach in the top-down setting of this problem would first detect all humans and then reconstruct each one of them independently. However, this type of prediction suffers from incohere…

Cited by 208PDFcodeScholar
2020

Monocular Expressive Body Regression through Body-Driven Attention

ECCV 2020poster

To understand how people look, interact, or perform tasks, we need to quickly and accurately capture their 3D body, face, and hands together from an RGB image. Most existing methods focus only on parts of the body. A few recent approaches reconstruct full expressive 3D humans from images using 3D bo…

2020

Reactive Semantic Planning in Unexplored Semantic Environments Using Deep Perceptual Feedback

RA-L 2020

This letter presents a reactive planning system that enriches the topological representation of an environment with a tightly integrated semantic representation, achieved by incorporating and exploiting advances in deep perceptual learning and probabilistic semantic reasoning. Our architecture combi

Cited by 34SourceScholar
2019

Convolutional Mesh Regression for Single-Image Human Shape Reconstruction

CVPR 2019oral

This paper addresses the problem of 3D human pose and shape estimation from a single image. Previous approaches consider a parametric model of the human body, SMPL, and attempt to regress the model parameters that give rise to a mesh consistent with image evidence. This parameter regression has been…

Cited by 670PDFScholar
2019

Expressive Body Capture: 3D Hands, Face, and Body From a Single Image

CVPR 2019oral

To facilitate the analysis of human actions, interactions and emotions, we compute a 3D model of human body pose, hand pose, and facial expression from a single monocular image. To achieve this, we use thousands of 3D scans to train a new, unified, 3D model of the human body, SMPL-X, that extends SM…

Cited by 2080PDFcodeScholar
2019

Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the Loop

ICCV 2019poster

Model-based human pose estimation is currently approached through two different paradigms. Optimization-based methods fit a parametric body model to 2D observations in an iterative manner, leading to accurate image-model alignments, but are often slow and sensitive to the initialization. In contrast…

Cited by 1232PDFScholar
2019

TexturePose: Supervising Human Mesh Estimation With Texture Consistency

ICCV 2019poster

This work addresses the problem of model-based human pose estimation. Recent approaches have made significant progress towards regressing the parameters of parametric human body models directly from images. Because of the absence of images with 3D shape ground truth, relevant approaches rely on 2D a…

Cited by 131PDFScholar
2018

Learning to Estimate 3D Human Pose and Shape From a Single Color Image

CVPR 2018poster

This work addresses the problem of estimating the full body 3D human pose and shape from a single color image. This is a task where iterative optimization-based solutions have typically prevailed, while Convolutional Networks (ConvNets) have suffered because of the lack of training data and their lo…

Cited by 785SourcePDFScholar
2017

6-DoF object pose from semantic keypoints

ICRA 2017poster

This paper presents a novel approach to estimating the continuous six degree of freedom (6-DoF) pose (3D translation and rotation) of an object from a single RGB image. The approach combines semantic keypoints predicted by a convolutional network (convnet) with a deformable shape model. Unlike prior…

Cited by 540SourceScholar
2017

Coarse-To-Fine Volumetric Prediction for Single-Image 3D Human Pose

CVPR 2017spotlight

This paper addresses the challenge of 3D human pose estimation from a single color image. Despite the general success of the end-to-end learning paradigm, top performing approaches employ a two-step solution consisting of a Convolutional Network (ConvNet) for 2D joint localization and a subsequent o…

Cited by 1177PDFScholar
2017

Harvesting Multiple Views for Marker-Less 3D Human Pose Annotations

CVPR 2017spotlight

Recent advances with Convolutional Networks (ConvNets) have shifted the bottleneck for many computer vision tasks to annotated data collection. In this paper, we present a geometry-driven approach to automatically collect annotations for human pose prediction tasks. Starting from a generic ConvNet f…

Cited by 248PDFScholar