← Search

Weipeng Xu

21 accepted papers

2025

REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning

CVPR 2025poster

We present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal…

Cited by 0SourcePDFScholar
2024

3D Hand Sequence Recovery from Real Blurry Images and Event Stream

ECCV 2024poster

"Although hands frequently exhibit motion blur due to their dynamic nature, existing approaches for 3D hand recovery often disregard the impact of motion blur in hand images. Blurry hand images contain hands from multiple time steps, lack precise hand location at a specific time step, and introduce…

Cited by 2SourcePDFScholar
2024

Authentic Hand Avatar from a Phone Scan via Universal Hand Model

CVPR 2024poster

The authentic 3D hand avatar with every identifiable information such as hand shapes and textures is necessary for immersive experiences in AR/VR. In this paper we present a universal hand model (UHM) which 1) can universally represent high-fidelity 3D hand meshes of arbitrary identities (IDs) and 2…

Cited by 5SourcePDFScholar
2024

Omnigrasp: Grasping Diverse Objects with Simulated Humanoids

NeurIPS 2024poster

We present a method for controlling a simulated humanoid to grasp an object and move it to follow an object's trajectory. Due to the challenges in controlling a humanoid with dexterous hands, prior methods often use a disembodied hand and only consider vertical lifts or short trajectories. This limi…

Cited by 1SourcePDFScholar
2024

Real-Time Simulated Avatar from Head-Mounted Sensors

CVPR 2024highlight

We present SimXR a method for controlling a simulated avatar from information (headset pose and cameras) obtained from AR / VR headsets. Due to the challenging viewpoint of head-mounted cameras the human body is often clipped out of view making traditional image-based egocentric pose estimation chal…

Cited by 8SourcePDFScholar
2024

Universal Humanoid Motion Representations for Physics-Based Control

ICLR 2024spotlight

We present a universal motion representation that encompasses a comprehensive range of motor skills for physics-based humanoid control. Due to the high dimensionality of humanoids and the inherent difficulties in reinforcement learning, prior methods have focused on learning skill embeddings for a n…

Cited by 58SourcePDFScholar
2023

A Dataset of Relighted 3D Interacting Hands

NeurIPS 2023poster

The two-hand interaction is one of the most challenging signals to analyze due to the self-similarity, complicated articulations, and occlusions of hands. Although several datasets have been proposed for the two-hand interaction analysis, all of them do not achieve 1) diverse and realistic image app…

2023

Perpetual Humanoid Control for Real-time Simulated Avatars

ICCV 2023poster

We present a physics-based humanoid controller that achieves high-fidelity motion imitation and fault-tolerant behavior in the presence of noisy input (e.g. pose estimates from video or generated from language) and unexpected falls. Our controller scales up to learning ten thousand motion clips with…

Cited by 85PDFScholar
2023

Scene-Aware Egocentric 3D Human Pose Estimation

CVPR 2023poster

Egocentric 3D human pose estimation with a single head-mounted fisheye camera has recently attracted attention due to its numerous applications in virtual and augmented reality. Existing methods still struggle in challenging poses where the human body is highly occluded or is closely interacting wit…

2022

Estimating Egocentric 3D Human Pose in the Wild With External Weak Supervision

CVPR 2022poster

Egocentric 3D human pose estimation with a single fisheye camera has drawn a significant amount of attention recently. However, existing methods struggle with pose estimation from in-the-wild images, because they can only be trained on synthetic data due to the unavailability of large-scale in-the-w…

Cited by 37PDFScholar
2022

HULC: 3D HUman Motion Capture with Pose Manifold SampLing and Dense Contact Guidance

ECCV 2022poster

"Marker-less monocular 3D human motion capture (MoCap) with scene interactions is a challenging research topic relevant for extended reality, robotics and virtual avatar generation. Due to the inherent depth ambiguity of monocular settings, 3D motions captured with existing methods often contain sev…

Cited by 31SourcePDFScholar
2021

Estimating Egocentric 3D Human Pose in Global Space

ICCV 2021poster

Egocentric 3D human pose estimation using a single fisheye camera has become popular recently as it allows capturing a wide range of daily activities in unconstrained environments, which is difficult for traditional outside-in motion capture with external cameras. However, existing methods have seve…

Cited by 83PDFcodeScholar
2021

Revitalizing Optimization for 3D Human Pose and Shape Estimation: A Sparse Constrained Formulation

ICCV 2021poster

We propose a novel sparse constrained formulation and from it derive a real-time optimization method for 3D human pose and shape estimation. Our optimization method, SCOPE (Sparse Constrained Optimization for 3D human Pose and shapE estimation), is orders of magnitude faster (avg. 4 ms convergence)…

Cited by 28PDFScholar
2020

DeepCap: Monocular Human Performance Capture Using Weak Supervision

CVPR 2020oral

Human performance capture is a highly important computer vision problem with many applications in movie production and virtual/augmented reality. Many previous performance capture approaches either required expensive multi-view setups or did not recover dense space-time coherent geometry with frame-…

Cited by 266PDFScholar
2020

EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event Camera

CVPR 2020oral

The high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement, the high data bandwidth and the consequent high computation overhead. In this paper, we propose EventCap -- the first…

Cited by 124PDFScholar
2020

Monocular Real-Time Hand Shape and Motion Capture Using Multi-Modal Data

CVPR 2020poster

We present a novel method for monocular hand shape and pose estimation at unprecedented runtime performance of 100fps and at state-of-the-art accuracy. This is enabled by a new learning based architecture designed such that it can make use of all the sources of available hand training data: image da…

Cited by 257PDFcodeScholar
2020

Neural Re-Rendering of Humans from a Single Image

ECCV 2020poster

Human re-rendering from a single image is a starkly under-constrained problem and state-of-the-art algorithms often exhibit un-desired artefacts, such as oversmoothing, unrealistic distortions of thebody parts and garments, or implausible changes of the texture. To ad-dress these challenges, we prop…

Cited by 91SourcePDFScholar
2019

In the Wild Human Pose Estimation Using Explicit 2D Features and Intermediate 3D Representations

CVPR 2019oral

Convolutional Neural Network based approaches for monocular 3D human pose estimation usually require a large amount of training images with 3D pose annotations. While it is feasible to provide 2D joint annotations for large corpora of in-the-wild images with humans, providing accurate 3D annotations…

Cited by 178PDFScholar
2018

A Hybrid Model for Identity Obfuscation by Face Replacement

ECCV 2018poster

As more and more personal photos are shared and tagged in social media, avoiding privacy risks such as unintended recognition, becomes increasingly challenging. We propose a new hybrid approach to obfuscate identities in photos by head replacement. Our approach combines state of the art parametric f…

Cited by 143SourcePDFScholar
2018

Video Based Reconstruction of 3D People Models

CVPR 2018poster

This paper describes how to obtain accurate 3D body models and texture of arbitrary people from a single, monocular video in which a person is moving. Based on a parametric body model, we present a robust processing pipeline achieving 3D model fits with 5mm accuracy also for clothed people. Our main…

2015

Deformable 3D Fusion: From Partial Dynamic 3D Observations to Complete 4D Models

ICCV 2015poster

Capturing the 3D motion of dynamic, non-rigid objects has attracted significant attention in computer vision. Existing methods typically require either complete 3D volumetric observations, or a shape template. In this paper, we introduce a template-less 4D reconstruction method that incrementally fu…

Cited by 16PDFScholar