← Search

Muhammed Kocabas

13 accepted papers

2025

BEDLAM2.0: Synthetic humans and cameras in motion

NeurIPS 2025oral

Inferring 3D human motion from video remains a challenging problem with many applications. While traditional methods estimate the human in image coordinates, many applications require human motion to be estimated in world coordinates. This is particularly challenging when there is both human and cam…

Cited by 0SourceScholar
2025

PromptHMR: Promptable Human Mesh Recovery

CVPR 2025poster

Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxiliary "side information" that could enhance reconstruction accuracy in such challe…

Cited by 0SourcePDFScholar
2024

HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video

CVPR 2024highlight

Since humans interact with diverse objects every day the holistic 3D capture of these interactions is important to understand and model human behaviour. However most existing methods for hand-object reconstruction from RGB either assume pre-scanned object templates or heavily rely on limited 3D hand…

2024

HUGS: Human Gaussian Splats

CVPR 2024poster

Recent advances in neural rendering have improved both training and rendering times by orders of magnitude. While these methods demonstrate state-of-the-art quality and speed they are designed for photogrammetry of static scenes and do not generalize well to freely moving humans in the environment.…

2023

ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

CVPR 2023poster

Humans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because there exist no datasets with ground-truth 3D annotations for…

2022

D-Grasp: Physically Plausible Dynamic Grasp Synthesis for Hand-Object Interactions

CVPR 2022poster

We introduce the dynamic grasp synthesis task: given an object with a known 6D pose and a grasp reference, our goal is to generate motions that move the object to a target 6D pose. This is challenging, because it requires reasoning about the complex articulation of the human hand and the intricate p…

Cited by 111PDFcodeScholar
2022

Human-Aware Object Placement for Visual Environment Reconstruction

CVPR 2022poster

Humans are in constant contact with the world as they move through it and interact with it. This contact is a vital source of information for understanding 3D humans, 3D scenes, and the interactions between them. In fact, we demonstrate that these human-scene interactions (HSIs) can be leveraged to…

Cited by 70PDFcodeScholar
2021

Learning To Regress Bodies From Images Using Differentiable Semantic Rendering

ICCV 2021poster

Learning to regress 3D human body shape and pose (e.g. SMPL parameters) from monocular images typically exploits losses on 2D keypoints, silhouettes, and/or part-segmentation when 3D training data is not available. Such losses, however, are limited because 2D keypoints do not supervise body shape an…

Cited by 65PDFcodeScholar
2021

PARE: Part Attention Regressor for 3D Human Body Estimation

ICCV 2021poster

Despite significant progress, we show that state of the art 3D human pose and shape estimation methods remain sensitive to partial occlusion and can produce dramatically wrong predictions although much of the body is observable. To address this, we introduce a soft attention mechanism, called the Pa…

Cited by 480PDFcodeScholar
2021

SPEC: Seeing People in the Wild With an Estimated Camera

ICCV 2021poster

Due to the lack of camera parameter information for in-the-wild images, existing 3D human pose and shape (HPS) estimation methods make several simplifying assumptions: weak-perspective projection, large constant focal length, and zero camera rotation. These assumptions often do not hold and we show,…

Cited by 160PDFcodeScholar
2020

VIBE: Video Inference for Human Body Pose and Shape Estimation

CVPR 2020poster

Human motion is fundamental to understanding behavior. Despite progress on single-image 3D pose and shape estimation, existing video-based state-of-the-art methods fail to produce accurate and natural motion sequences due to a lack of ground-truth 3D motion data for training. To address this problem…

Cited by 1216PDFcodeScholar
2018

MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual Network

ECCV 2018poster

In this paper, we present MultiPoseNet, a novel bottom-up multi-person pose estimation architecture that combines a multi-task model with a novel assignment method. MultiPoseNet can jointly handle person detection, person segmentation and pose estimation problems. The novel assignment method is impl…