← Search

Chengcheng Tang

16 accepted papers

2026

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens

CVPR 2026

Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language generation and understanding remains largely underexplored. Existing approaches often fine-tune large language models (LLM

Cited by 0SourcecodeScholar
2025

EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba

ICCV 2025poster

Estimating human dance motion is a challenging task with various industrial applications. Recently, many efforts have focused on predicting human dance motion using either egocentric video or music as input. However, the task of jointly estimating human motion from both egocentric video and music re…

Cited by 0SourcePDFScholar
2025

HuMoCon: Concept Discovery for Human Motion Understanding

CVPR 2025poster

We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract semantically meaningful and generalizable features. HuMoCon add…

Cited by 0SourcePDFScholar
2025

PHD: Personalized 3D Human Body Fitting with Point Diffusion

ICCV 2025poster

We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these met…

2024

CigTime: Corrective Instruction Generation Through Inverse Motion Editing

NeurIPS 2024poster

Recent advancements in models linking natural language with human motions have shown significant promise in motion generation and editing based on instructional text. Motivated by applications in sports coaching and motor skill learning, we investigate the inverse problem: generating corrective inst…

Cited by 0SourcePDFScholar
2024

EgoBody3M: Egocentric Body Tracking on a VR Headset using a Diverse Dataset

ECCV 2024poster

"Accurate tracking of a user’s body pose while wearing a virtual reality (VR), augmented reality (AR) or mixed reality (MR) headset is a prerequisite for authentic self-expression, natural social presence, and intuitive user interfaces. Existing body tracking approaches on VR/AR devices are either u…

Cited by 4SourcePDFScholar
2023

EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild

ICCV 2023poster

We present EMDB, the Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. EMDB is a novel dataset that contains high-quality 3D SMPL pose and shape parameters with global body and camera trajectories for in-the-wild videos. We use body-worn, wireless electromagnetic (EM) sensors a…

Cited by 52PDFcodeScholar
2023

MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded Convolutions

ICCV 2023poster

Convolutional neural network inference on video input is computationally expensive and requires high memory bandwidth. Recently, DeltaCNN managed to reduce the cost by only processing pixels with significant updates over the previous frame. However, DeltaCNN relies on static camera input. Moving cam…

Cited by 5PDFScholar
2023

Social Diffusion: Long-term Multiple Human Motion Anticipation

ICCV 2023poster

We propose Social Diffusion, a novel method for short-term and long-term forecasting of the motion of multiple persons as well as their social interactions. Jointly forecasting motions for multiple persons involved in social activities is inherently a challenging problem due to the interdependenci…

Cited by 20PDFcodeScholar
2022

DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in Videos

CVPR 2022poster

Convolutional neural network inference on video data requires powerful hardware for real-time processing. Given the inherent coherence across consecutive frames, large parts of a video typically change little. By skipping identical image regions and truncating insignificant pixel updates, computatio…

Cited by 34PDFcodeScholar
2022

Multiview Human Body Reconstruction from Uncalibrated Cameras

NeurIPS 2022accept

We present a new method to reconstruct 3D human body pose and shape by fusing visual features from multiview images captured by uncalibrated cameras. Existing multiview approaches often use spatial camera calibration (intrinsic and extrinsic parameters) to geometrically align and fuse visual feature…

Cited by 21SourcePDFScholar
2022

PressureVision: Estimating Hand Pressure from a Single RGB Image

ECCV 2022poster

"People often interact with their surroundings by applying pressure with their hands. While hand pressure can be measured by placing pressure sensors between the hand and the environment, doing so can alter contact mechanics, interfere with human tactile perception, require costly sensors, and scale…

2022

Visual Pressure Estimation and Control for Soft Robotic Grippers

IROS 2022poster

Soft robotic grippers facilitate contact-rich manipulation, including robust grasping of varied objects. Yet the beneficial compliance of a soft gripper also results in significant deformation that can make precision manipulation challenging. We present visual pressure estimation & control (VPEC), a…

Cited by 6SourcecodeScholar
2021

ContactOpt: Optimizing Contact To Improve Grasps

CVPR 2021poster

Physical contact between hands and objects plays a critical role in human grasps. We show that optimizing the pose of a hand to achieve expected contact with an object can improve hand poses inferred via image-based methods. Given a hand mesh and an object mesh, a deep model trained on ground truth…

Cited by 147PDFcodeScholar
2021

EM-POSE: 3D Human Pose Estimation From Sparse Electromagnetic Trackers

ICCV 2021poster

Fully immersive experiences in AR/VR depend on reconstructing the full body pose of the user without restricting their motion. In this paper we study the use of body-worn electromagnetic (EM) field-based sensing for the task of 3D human pose reconstruction. To this end, we present a method to estima…

Cited by 40PDFcodeScholar
2020

ContactPose: A Dataset of Grasps with Object Contact and Hand Pose

ECCV 2020poster

Grasping is natural for humans. However, it involves complex hand configurations and soft tissue deformation that can result in complicated regions of contact between the hand and the object. Understanding and modeling this contact can potentially improve hand models, AR/VR experiences, and robotic…

Cited by 234SourcePDFScholar