← Search

Jiefeng Li

22 accepted papers

2026

CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

CVPR 2026

Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming, and robot learning. However, inferring 4D interactions from a single RGB view is highly challenging due to the unknown object and human information, dep

Cited by 0SourcecodeScholar
2025

BLADE: Single-view Body Mesh Estimation through Accurate Depth Estimation

CVPR 2025poster

Single-image human mesh recovery is a challenging task due to the ill-posed nature of simultaneous body shape, pose, and camera estimation. Existing estimators work well on images taken from afar, but they break down as the person moves close to the camera. Moreover, current methods fail to achieve…

Cited by 0SourcePDFScholar
2025

GENMO: A GENeralist Model for Human MOtion

ICCV 2025poster

Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes, while motion estimation models aim to reconstruct accurate mot…

Cited by 0SourcePDFScholar
2024

Denoising Vision Transformers

ECCV 2024oral

"We study a crucial yet often overlooked issue inherent to Vision Transformers (ViTs): feature maps of these models exhibit grid-like artifacts (“Original features” in fig:teaser), which hurt the performance of ViTs in downstream dense prediction tasks such as semantic segmentation, depth prediction…

2024

ShapeBoost: Boosting Human Shape Estimation with Part-Based Parameterization and Clothing-Preserving Augmentation

AAAI 2024technical

Accurate human shape recovery from a monocular RGB image is a challenging task because humans come in different shapes and sizes and wear different clothes. In this paper, we propose ShapeBoost, a new human shape recovery framework that achieves pixel-level alignment even for rare body shapes and hi…

Cited by 2SourcePDFScholar
2023

Learning Analytical Posterior Probability for Human Mesh Recovery

CVPR 2023poster

Despite various probabilistic methods for modeling the uncertainty and ambiguity in human mesh recovery, their overall precision is limited because existing formulations for joint rotations are either not constrained to SO(3) or difficult to learn for neural networks. To address such an issue, we de…

2023

NIKI: Neural Inverse Kinematics With Invertible Neural Networks for 3D Human Pose and Shape Estimation

CVPR 2023poster

With the progress of 3D human pose and shape estimation, state-of-the-art methods can either be robust to occlusions or obtain pixel-aligned accuracy in non-occlusion cases. However, they cannot obtain robustness and mesh-image alignment at the same time. In this work, we present NIKI (Neural Invers…

2022

ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis

CVPR 2022oral

Estimating the articulated 3D hand-object pose from a single RGB image is a highly ambiguous and challenging problem, requiring large-scale datasets that contain diverse hand poses, object types, and camera viewpoints. Most real-world datasets lack these diversities. In contrast, data synthesis can…

Cited by 98PDFcodeScholar
2022

Constructing Balance from Imbalance for Long-Tailed Image Recognition

ECCV 2022poster

"Long-tailed image recognition presents massive challenges to deep learning systems since the imbalance between majority (head) classes and minority (tail) classes severely skews the data-driven deep neural networks. Previous methods tackle with data imbalance from the viewpoints of data distributio…

2022

Correlation Field for Boosting 3D Object Detection in Structured Scenes

AAAI 2022technical

Data augmentation is an efficient way to elevate 3D object detection performance. In this paper, we propose a simple but effective online crop-and-paste data augmentation pipeline for structured 3D point cloud scenes, named CorrelaBoost. Observing that 3D objects should have reasonable relative posi…

Cited by 10SourcePDFScholar
2022

D&D: Learning Human Dynamics from Dynamic Camera

ECCV 2022poster

"3D human pose estimation from a monocular video has recently seen significant improvements. However, most state-of-the-art methods are kinematics-based, which are prone to physically implausible motions with pronounced artifacts. Current dynamics-based methods can predict physically plausible motio…

2022

SmoothNet: A Plug-and-Play Network for Refining Human Poses in Videos

ECCV 2022poster

"When analyzing human motion videos, the output jitters from existing pose estimators are highly-unbalanced with varied estimation errors across frames. Most frames in a video are relatively easy to estimate and only suffer from slight jitters. In contrast, for rarely seen or occluded actions, the e…

2022

Unified and Fast Human Trajectory Prediction Via Conditionally Parameterized Normalizing Flow

RA-L 2022

Human trajectory prediction is crucial for service robots, autonomous driving and advanced driver assistant systems. Current top-performing methods mainly rely on intractable generative models to learn a distribution of future trajectories, and sample multiple plausible ones as prediction results. I

Cited by 14SourceScholar
2022

Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive Learning

AAAI 2022technical

We study the unsupervised representation learning for the semantic segmentation task. Different from previous works that aim at providing unsupervised pre-trained backbones for segmentation models which need further supervised fine-tune, here, we focus on providing representation that is only traine…

Cited by 11SourcePDFScholar
2021

CPF: Learning a Contact Potential Field To Model the Hand-Object Interaction

ICCV 2021poster

Modeling the hand-object (HO) interaction not only requires estimation of the HO pose, but also pays attention to the contact due to their interaction. Significant progress has been made in estimating hand and object separately with deep learning methods, simultaneous HO pose estimation and contact…

Cited by 141PDFcodeScholar
2021

Human Pose Regression With Residual Log-Likelihood Estimation

ICCV 2021poster

Heatmap-based methods dominate in the field of human pose estimation by modelling the output distribution through likelihood heatmaps. In contrast, regression-based methods are more efficient but suffer from inferior performance. In this work, we explore maximum likelihood estimation (MLE) to develo…

Cited by 277PDFcodeScholar
2021

HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape Estimation

CVPR 2021poster

Model-based 3D pose and shape estimation methods reconstruct a full 3D mesh for the human body by estimating several parameters. However, learning the abstract parameters is a highly non-linear process and suffers from image-model misalignment, leading to mediocre model performance. In contrast, 3D…

Cited by 469PDFcodeScholar
2021

TDAF: Top-Down Attention Framework for Vision Tasks

AAAI 2021technical

Human attention mechanisms often work in a top-down manner, yet it is not well explored in vision research. Here, we propose the Top-Down Attention Framework (TDAF) to capture top-down attentions, which can be easily adopted in most existing models. The designed Recursive Dual-Directional Nested Str…

Cited by 13SourcePDFScholar
2020

Detailed 2D-3D Joint Representation for Human-Object Interaction

CVPR 2020poster

Human-Object Interaction (HOI) detection lies at the core of action understanding. Besides 2D information such as human/object appearance and locations, 3D pose is also usually utilized in HOI learning since its view-independence. However, rough 3D body joints just carry sparse body information and…

Cited by 175PDFcodeScholar
2020

HMOR: Hierarchical Multi-Person Ordinal Relations for Monocular Multi-Person 3D Pose Estimation

ECCV 2020poster

Remarkable progress has been made in 3D human pose estimation from a monocular RGB camera. However, only a few studies explored 3D multi-person cases. In this paper, we attempt to address the lack of a global perspective of the top-down approaches by introducing a novel form of supervision - Hierarc…

Cited by 75SourcePDFScholar
2019

CrowdPose: Efficient Crowded Scenes Pose Estimation and a New Benchmark

CVPR 2019oral

Multi-person pose estimation is fundamental to many computer vision tasks and has made significant progress in recent years. However, few previous methods explored the problem of pose estimation in crowded scenes while it remains challenging and inevitable in many scenarios. Moreover, current benchm…

Cited by 699PDFcodeScholar