← Search

Lan Xu

54 accepted papers

2026

FlashCap: Millisecond-Accurate Human Motion Capture via Flashing LEDs and Event-Based Vision

CVPR 2026

Precise motion timing (PMT) is crucial for swift motion analysis. A millisecond difference may determine victory or defeat in sports competitions. Despite substantial progress in human pose estimation (HPE), PMT remains largely overlooked by the HPE community due to the limited availability of high-

Cited by 0SourceScholar
2026

Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy

ICLR 2026poster

Human-object interaction (HOI) synthesis is crucial for applications in animation, simulation, and robotics. However, existing approaches either rely on expensive motion capture data or require manual reward engineering, limiting their scalability and generalizability. In this work, we introduce the…

Cited by 0SourcecodeScholar
2026

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

CVPR 2026

Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAg

Cited by 0SourcecodeScholar
2026

Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects

ICRA 2026poster

A deep understanding of kinematic structures is essential for robot motion and interaction with the environment. Such understanding is captured through articulated objects, which are essential for physical simulation, motion planning, and policy learning. However, creating these models, particularly…

2026

MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations

ICML 2026poster

Human motion inherently exhibits a sophisticated temporal hierarchical architecture, spanning from global low-frequency trajectories to local high-frequency dynamics. Inspired by this intrinsic property and the success of multi-scale autoregressive modeling in vision, we propose MotionMAR, a novel f…

Cited by 0SourceScholar
2026

Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid Robots

CVPR 2026

Humanoid robots have achieved significant progress in motion generation and control, exhibiting movements that appear increasingly natural and human-like. Inspired by the Turing Test, we propose the Motion Turing Test, a framework that evaluates whether human observers can discriminate between human

Cited by 0SourceScholar
2026

WildCap: Facial Albedo Capture in the Wild via Hybrid Inverse Rendering

CVPR 2026

Existing methods achieve high-quality facial albedo capture under controllable lighting, which increases capture cost and limits usability. We propose WildCap, a novel method for high-quality facial albedo capture from a smartphone video recorded in the wild. To disentangle high-quality albedo from

Cited by 0SourcecodeScholar
2025

4DGCPro: Efficient Hierarchical 4D Gaussian Compression for Progressive Volumetric Video Streaming

NeurIPS 2025poster

Achieving seamless viewing of high-fidelity volumetric video, comparable to 2D video experiences, remains an open challenge. Existing volumetric video compression methods either lack the flexibility to adjust quality and bitrate within a single model for efficient streaming across diverse networks a…

Cited by 0SourceScholar
2025

Capturing the Unseen: Vision-Free Facial Motion Capture Using Inertial Measurement Units

AAAI 2025technical

We present Capturing the Unseen (CAPUS), a novel facial motion capture (MoCap) technique that operates without visual signals. CAPUS leverages miniaturized Inertial Measurement Units (IMUs) as a new sensing modality for facial motion capture. While IMUs have become essential in full-body MoCap for t…

Cited by 0SourcePDFScholar
2025

ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World Coordinate

CVPR 2025highlight

Human Motion Recovery (HMR) research mainly focuses on ground-based motions such as running. The study on capturing climbing motion, an off-ground motion, is sparse. This is partly due to the limited availability of climbing motion datasets, especially large-scale and challenging 3D labeled datasets…

Cited by 0SourcePDFScholar
2025

PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding

NeurIPS 2025poster

Understanding objects at the level of their constituent parts is fundamental to advancing computer vision, graphics, and robotics. While datasets like PartNet have driven progress in 3D part understanding, their reliance on untextured geometries and expert-dependent annotation limits scalability and…

Cited by 0SourceScholar
2025

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

CVPR 2025poster

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform general dynamic scenes. In this paper, we present RePerformer, a novel Gaussian-based rep…

Cited by 0SourcePDFScholar
2025

SCOPE: Sign Language Contextual Processing with Embedding from LLMs

AAAI 2025technical

Sign languages, used by around 70 million Deaf individuals globally, are visual languages that convey visual and contextual information. Current methods in vision-based sign language recognition (SLR) and translation (SLT) struggle with dialogue scenes due to limited dataset diversity and the neglec…

2025

SMGDiff: Soccer Motion Generation using Diffusion Probabilistic Models

ICCV 2025poster

Soccer is a globally renowned sport with significant applications in video games and VR/AR. However, generating realistic soccer motions remains challenging due to the intricate interactions between the player and the ball. In this paper, we introduce SMGDiff, a novel two-stage framework for generat…

Cited by 0SourcePDFScholar
2025

Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis

ICCV 2025poster

Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate progress in kinematic motion generation, they often fail to address the fundamental tension between real-time responsivene…

Cited by 0SourcePDFScholar
2024

A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals

CVPR 2024poster

Estimating full-body human motion via sparse tracking signals from head-mounted displays and hand controllers in 3D scenes is crucial to applications in AR/VR. One of the biggest challenges to this task is the one-to-many mapping from sparse observations to dense full-body motions which endowed inhe…

2024

BOTH2Hands: Inferring 3D Hands from Both Text Prompts and Body Dynamics

CVPR 2024poster

The recently emerging text-to-motion advances have spired numerous attempts for convenient and interactive human motion generation. Yet existing methods are largely limited to generating body motions only without considering the rich two-hand motions let alone handling various conditions like body d…

2024

HOI-M^3: Capture Multiple Humans and Objects Interaction within Contextual Environment

CVPR 2024highlight

Humans naturally interact with both others and the surrounding multiple objects engaging in various social activities. However recent advances in modeling human-object interactions mostly focus on perceiving isolated individuals and objects due to fundamental data scarcity. In this paper we introduc…

Cited by 12SourcePDFScholar
2024

HiFi4G: High-Fidelity Human Performance Rendering via Compact Gaussian Splatting

CVPR 2024poster

We have recently seen tremendous progress in photo-real human modeling and rendering. Yet efficiently rendering realistic human performance and integrating it into the rasterization pipeline remains challenging. In this paper we present HiFi4G an explicit and compact Gaussian-based approach for high…

Cited by 49SourcePDFScholar
2024

HybridGait: A Benchmark for Spatial-Temporal Cloth-Changing Gait Recognition with Hybrid Explorations

AAAI 2024technical

Existing gait recognition benchmarks mostly include minor clothing variations in the laboratory environments, but lack persistent changes in appearance over time and space. In this paper, we propose the first in-the-wild benchmark CCGait for cloth-changing gait recognition, which incorporates divers…

2024

I'M HOI: Inertia-aware Monocular Capture of 3D Human-Object Interactions

CVPR 2024poster

We are living in a world surrounded by diverse and "smart" devices with rich modalities of sensing ability. Conveniently capturing the interactions between us humans and these objects remains far-reaching. In this paper we present I'm-HOI a monocular scheme to faithfully capture the 3D motions of bo…

Cited by 7SourcePDFScholar
2024

LiveHPS: LiDAR-based Scene-level Human Pose and Shape Estimation in Free Environment

CVPR 2024highlight

For human-centric large-scale scenes fine-grained modeling for 3D human global pose and shape is significant for scene understanding and can benefit many real-world applications. In this paper we present LiveHPS a novel single-LiDAR-based approach for scene-level human pose and shape estimation with…

Cited by 15SourcePDFScholar
2024

OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers

CVPR 2024poster

We have recently seen tremendous progress in realistic text-to-motion generation. Yet the existing methods often fail or produce implausible motions with unseen text inputs which limits the applications. In this paper we present OMG a novel framework which enables compelling motion generation from z…

2024

RELI11D: A Comprehensive Multimodal Human Motion Dataset and Method

CVPR 2024poster

Comprehensive capturing of human motions requires both accurate captures of complex poses and precise localization of the human within scenes. Most of the HPE datasets and methods primarily rely on RGB LiDAR or IMU data. However solely using these modalities or a combination of them may not be adequ…

Cited by 8SourcePDFScholar
2024

VideoRF: Rendering Dynamic Radiance Fields as 2D Feature Video Streams

CVPR 2024poster

Neural Radiance Fields (NeRFs) excel in photorealistically rendering static scenes. However rendering dynamic long-duration radiance fields on ubiquitous devices remains challenging due to data storage and computational constraints. In this paper we introduce VideoRF the first approach to enable rea…

2023

CIMI4D: A Large Multimodal Climbing Motion Dataset Under Human-Scene Interactions

CVPR 2023poster

Motion capture is a long-standing research problem. Although it has been studied for decades, the majority of research focus on ground-based movements such as walking, sitting, dancing, etc. Off-grounded actions such as climbing are largely overlooked. As an important type of action in sports and fi…

Cited by 30SourcePDFScholar
2023

Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator

NeurIPS 2023poster

Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt. This study focuses on zero-shot text-to-video generation considering the data- and cost-efficient. To generate a s…

2023

HumanGen: Generating Human Radiance Fields With Explicit Priors

CVPR 2023poster

Recent years have witnessed the tremendous progress of 3D GANs for generating view-consistent radiance fields with photo-realism. Yet, high-quality generation of human radiance fields remains challenging, partially due to the limited human-related priors adopted in existing methods. We present Human…

Cited by 38SourcePDFScholar
2023

HybridCap: Inertia-Aid Monocular Capture of Challenging Human Motions

AAAI 2023technical

Monocular 3D motion capture (mocap) is beneficial to many applications. The use of a single camera, however, often fails to handle occlusions of different body parts and hence it is limited to capture relatively simple movements. We present a light-weight, hybrid mocap technique called HybridCap tha…

2023

IKOL: Inverse Kinematics Optimization Layer for 3D Human Pose and Shape Estimation via Gauss-Newton Differentiation

AAAI 2023technical

This paper presents an inverse kinematic optimization layer (IKOL) for 3D human pose and shape estimation that leverages the strength of both optimization- and regression-based methods within an end-to-end framework. IKOL involves a nonconvex optimization that establishes an implicit mapping from an…

2023

Instant-NVR: Instant Neural Volumetric Rendering for Human-Object Interactions From Monocular RGBD Stream

CVPR 2023poster

Convenient 4D modeling of human-object interactions is essential for numerous applications. However, monocular tracking and rendering of complex interaction scenarios remain challenging. In this paper, we propose Instant-NVR, a neural approach for instant volumetric human-object tracking and renderi…

Cited by 21SourcePDFScholar
2023

Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos

CVPR 2023poster

The success of the Neural Radiance Fields (NeRFs) for modeling and free-view rendering static objects has inspired numerous attempts on dynamic scenes. Current techniques that utilize neural rendering for facilitating free-view videos (FVVs) are restricted to either offline rendering or are capable…

Cited by 66SourcePDFScholar
2023

NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions

CVPR 2023poster

Humans constantly interact with objects in daily life tasks. Capturing such processes and subsequently conducting visual inferences from a fixed viewpoint suffers from occlusions, shape and texture ambiguities, motions, etc. To mitigate the problem, it is essential to build a training dataset that c…

2023

Relightable Neural Human Assets From Multi-View Gradient Illuminations

CVPR 2023poster

Human modeling and relighting are two fundamental problems in computer vision and graphics, where high-quality datasets can largely facilitate related research. However, most existing human datasets only provide multi-view human images captured under the same illumination. Although valuable for mode…

2023

SLOPER4D: A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban Environments

CVPR 2023poster

We present SLOPER4D, a novel scene-aware dataset collected in large urban environments to facilitate the research of global human pose estimation (GHPE) with human-scene interaction in the wild. Employing a head-mounted device integrated with a LiDAR and camera, we record 12 human subjects' activiti…

2023

StackFLOW: Monocular Human-Object Reconstruction by Stacked Normalizing Flow with Offset

IJCAI 2023poster

Modeling and capturing the 3D spatial arrangement of the human and the object is the key to perceiving 3D human-object interaction from monocular images. In this work, we propose to use the Human-Object Offset between anchors which are densely sampled from the surface of human mesh and object mesh t…

2023

Weakly Supervised 3D Multi-Person Pose Estimation for Large-Scale Scenes Based on Monocular Camera and Single LiDAR

AAAI 2023technical

Depth estimation is usually ill-posed and ambiguous for monocular camera-based 3D multi-person pose estimation. Since LiDAR can capture accurate depth information in long-range scenes, it can benefit both the global localization of individuals and the 3D pose estimation by providing rich geometry fe…

2022

Anisotropic Fourier Features for Neural Image-Based Rendering and Relighting

AAAI 2022technical

Recent neural rendering techniques have greatly benefited image-based modeling and relighting tasks. They provide a continuous, compact, and parallelable representation by modeling the plenoptic function as multilayer perceptrons (MLPs). However, vanilla MLPs suffer from spectral biases on multidime…

Cited by 7SourcePDFScholar
2022

Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-Time

CVPR 2022oral

Implicit neural representations such as Neural Radiance Field (NeRF) have focused mainly on modeling static objects captured under multi-view settings where real-time rendering can be achieved with smart data structures, e.g., PlenOctree. In this paper, we present a novel Fourier PlenOctree (FPO) te…

Cited by 178PDFScholar
2022

HSC4D: Human-Centered 4D Scene Capture in Large-Scale Indoor-Outdoor Space Using Wearable IMUs and LiDAR

CVPR 2022poster

We propose Human-centered 4D Scene Capture (HSC4D) to accurately and efficiently create a dynamic digital world, containing large-scale indoor-outdoor scenes, diverse human motions, and rich interactions between humans and environments. Using only body-mounted IMUs and LiDAR, HSC4D is space-free wit…

Cited by 36PDFcodeScholar
2022

HumanNeRF: Efficiently Generated Human Radiance Field From Sparse Inputs

CVPR 2022poster

Recent neural human representations can produce high-quality multi-view rendering but require using dense multi-view inputs and costly training. They are hence largely limited to static models as training each frame is infeasible. We present HumanNeRF - a neural representation with efficient general…

Cited by 227PDFScholar
2022

LiDARCap: Long-Range Marker-Less 3D Human Motion Capture With LiDAR Point Clouds

CVPR 2022poster

Existing motion capture datasets are largely short-range and cannot yet fit the need of long-range applications. We propose LiDARHuman26M, a new human motion capture dataset captured by LiDAR at a much longer range to overcome this limitation. Our dataset also includes the ground truth human motions…

Cited by 62PDFScholar
2022

NeuralHOFusion: Neural Volumetric Rendering Under Human-Object Interactions

CVPR 2022poster

4D modeling of human-object interactions is critical for numerous applications. However, efficient volumetric capture and rendering of complex interaction scenarios, especially from sparse inputs, remain challenging. In this paper, we propose NeuralHOFusion, a neural approach for volumetric human-ob…

Cited by 50PDFScholar
2022

STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded Scenes

CVPR 2022poster

Accurately detecting and tracking pedestrians in 3D space is challenging due to large variations in rotations, poses and scales. The situation becomes even worse for dense crowds with severe occlusions. However, existing benchmarks either only provide 2D annotations, or have limited 3D annotations w…

Cited by 49PDFcodeScholar
2021

ChallenCap: Monocular 3D Capture of Challenging Human Performances Using Multi-Modal References

CVPR 2021poster

Capturing challenging human motions is critical for numerous applications, but it suffers from complex motion patterns and severe self-occlusion under the monocular setting. In this paper, we propose ChallenCap --- a template-based approach to capture challenging 3D human motions using a single RGB…

Cited by 28PDFScholar
2021

Few-shot Neural Human Performance Rendering from Sparse RGBD Videos

IJCAI 2021poster

Recent neural rendering approaches for human activities achieve remarkable view synthesis results, but still rely on dense input views or dense training with all the capture frames, leading to deployment difficulty and inefficient training overload. However, existing advances will be ill-posed if th…

Cited by 17SourcePDFScholar
2021

GNeRF: GAN-Based Neural Radiance Field Without Posed Camera

ICCV 2021poster

We introduce GNeRF, a framework to marry Generative Adversarial Networks (GAN) with Neural Radiance Field (NeRF) reconstruction for the complex scenarios with unknown and even randomly initialized camera poses. Recent NeRF-based advances have gained popularity for remarkable realistic novel view syn…

Cited by 223PDFcodeScholar
2021

Neural Video Portrait Relighting in Real-Time via Consistency Modeling

ICCV 2021poster

Video portraits relighting is critical in user-facing human photography, especially for immersive VR/AR experience. Recent advances still fail to recover consistent relit result under dynamic illuminations from monocular RGB stream, suffering from the lack of video consistency supervision. In this p…

Cited by 47PDFcodeScholar
2021

NeuralHumanFVV: Real-Time Neural Volumetric Human Performance Rendering Using RGB Cameras

CVPR 2021poster

4D reconstruction and rendering of human activities is critical for immersive VR/AR experience. Recent advances still fail to recover fine geometry and texture results with the level of detail present in the input images from sparse multi-view RGB cameras. In this paper, we propose NeuralHumanFVV, a…

Cited by 50PDFScholar
2021

PIANO: A Parametric Hand Bone Model from Magnetic Resonance Imaging

IJCAI 2021poster

Hand modeling is critical for immersive VR/AR, action understanding, or human healthcare. Existing parametric models account only for hand shape, pose, or texture, without modeling the anatomical attributes like bone, which is essential for realistic hand biomechanics analysis. In this paper, we pre…

2020

EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event Camera

CVPR 2020oral

The high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement, the high data bandwidth and the consequent high computation overhead. In this paper, we propose EventCap -- the first…

Cited by 124PDFScholar
2020

RobustFusion: Human Volumetric Capture with Data-driven Visual Cues using a RGBD Camera

ECCV 2020poster

High-quality and complete 4D reconstruction of human activities is critical for immersive VR/AR experience, but it suffers from inherent self-scanning constraint and consequent fragile tracking under the monocular setting. In this paper, inspired by the huge potential of learning-based human modelin…

Cited by 106SourcePDFScholar