← Search

Sai Kumar Dwivedi

9 accepted papers

2026

EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR

CVPR 2026

Egocentric 3D human motion estimation is essential for AR/VR experiences, yet remains challenging due to limited body coverage from the egocentric viewpoint, frequent occlusions, and scarce labeled data. We present EgoPoseFormer v2, a method that addresses these challenges through two key contributi

Cited by 0SourceScholar
2025

InteractVLM: 3D Interaction Reasoning from 2D Foundational Models

CVPR 2025poster

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth ambiguities, and widely varying object shapes. Existing methods r…

2025

PICO: Reconstructing 3D People In Contact with Objects

CVPR 2025poster

Recovering 3D Human-Object Interaction (HOI) from single color images is challenging due to depth ambiguities, occlusions, and the huge variation in object shape and appearance. Thus, past work requires controlled settings such as known object shapes and contacts, and tackles only limited object cla…

Cited by 1SourcePDFScholar
2025

SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single Image

ICCV 2025poster

Recovering 3D object pose and shape from a single image is a challenging and ill-posed problem. This is due to strong (self-)occlusions, depth ambiguities, the vast intra- and inter-class shape variance, and the lack of 3D ground truth for natural images. Existing deep-network methods are trained on…

2024

ChatPose: Chatting about 3D Human Pose

CVPR 2024poster

We introduce ChatPose a framework employing Large Language Models (LLMs) to understand and reason about 3D human poses from images or textual descriptions. Our work is motivated by the human ability to intuitively understand postures from a single image or a brief description a process that intertwi…

2024

TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

CVPR 2024poster

We address the problem of regressing 3D human pose and shape from a single image with a focus on 3D accuracy. The current best methods leverage large datasets of 3D pseudo-ground-truth (p-GT) and 2D keypoints leading to robust performance. With such methods however we observe a paradoxical decline i…

2023

Detecting Human-Object Contact in Images

CVPR 2023poster

Humans constantly contact objects to move and perform tasks. Thus, detecting human-object contact is important for building human-centered artificial intelligence. However, there exists no robust method to detect contact between the body and the scene from an image, and there exists no dataset to le…

2021

Learning To Regress Bodies From Images Using Differentiable Semantic Rendering

ICCV 2021poster

Learning to regress 3D human body shape and pose (e.g. SMPL parameters) from monocular images typically exploits losses on 2D keypoints, silhouettes, and/or part-segmentation when 3D training data is not available. Such losses, however, are limited because 2D keypoints do not supervise body shape an…

Cited by 65PDFcodeScholar
2019

Out-Of-Distribution Detection for Generalized Zero-Shot Action Recognition

CVPR 2019poster

Generalized zero-shot action recognition is a challenging problem, where the task is to recognize new action categories that are unavailable during the training stage, in addition to the seen action categories. Existing approaches suffer from the inherent bias of the learned classifier towards the s…

Cited by 193PDFcodeScholar