← Search

Dimitrios Tzionas

23 accepted papers

2026

RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos

CVPR 2026

Reconstructing people, objects, and their interactions in 3D is a long-standing and fundamental goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each other, and camera and object motion entangle

Cited by 0SourcecodeScholar
2025

InterDyn: Controllable Interactive Dynamics with Video Diffusion Models

CVPR 2025poster

Predicting the dynamics of interacting objects is essential for both humans and intelligent systems. However, existing approaches are limited to simplified, toy settings and lack generalizability to complex, real-world environments. Recent advances in generative models have enabled the prediction of…

Cited by 2SourcePDFScholar
2025

InteractVLM: 3D Interaction Reasoning from 2D Foundational Models

CVPR 2025poster

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth ambiguities, and widely varying object shapes. Existing methods r…

2025

PICO: Reconstructing 3D People In Contact with Objects

CVPR 2025poster

Recovering 3D Human-Object Interaction (HOI) from single color images is challenging due to depth ambiguities, occlusions, and the huge variation in object shape and appearance. Thus, past work requires controlled settings such as known object shapes and contacts, and tackles only limited object cla…

Cited by 1SourcePDFScholar
2025

SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single Image

ICCV 2025poster

Recovering 3D object pose and shape from a single image is a challenging and ill-posed problem. This is due to strong (self-)occlusions, depth ambiguities, the vast intra- and inter-class shape variance, and the lack of 3D ground truth for natural images. Existing deep-network methods are trained on…

2023

3D Human Pose Estimation via Intuitive Physics

CVPR 2023poster

Estimating 3D humans from images often produces implausible bodies that lean, float, or penetrate the floor. Such methods ignore the fact that bodies are typically supported by the scene. A physics engine can be used to enforce physical plausibility, but these are not differentiable, rely on unreali…

Cited by 98SourcePDFScholar
2023

ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

CVPR 2023poster

Humans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because there exist no datasets with ground-truth 3D annotations for…

2023

DECO: Dense Estimation of 3D Human-Scene Contact In The Wild

ICCV 2023oral

Understanding how humans use physical contact to interact with the world is key to enabling human-centric artificial intelligence. While inferring 3D contact is crucial for modeling realistic and physically-plausible human-object interactions, existing methods either focus on 2D, consider body joint…

Cited by 25PDFcodeScholar
2023

Detecting Human-Object Contact in Images

CVPR 2023poster

Humans constantly contact objects to move and perform tasks. Thus, detecting human-object contact is important for building human-centered artificial intelligence. However, there exists no robust method to detect contact between the body and the scene from an image, and there exists no dataset to le…

2023

ECON: Explicit Clothed Humans Optimized via Normal Integration

CVPR 2023highlight

The combination of deep learning, artist-curated scans, and Implicit Functions (IF), is enabling the creation of detailed, clothed, 3D humans from images. However, existing methods are far from perfect. IF-based methods recover free-form geometry, but produce disembodied limbs or degenerate shapes f…

2023

Reconstructing Signing Avatars From Video Using Linguistic Priors

CVPR 2023poster

Sign language (SL) is the primary method of communication for the 70 million Deaf people around the world. Video dictionaries of isolated signs are a core SL learning tool. Replacing these with 3D avatars can aid learning and enable AR/VR applications, improving access to technology and online media…

Cited by 15SourcePDFScholar
2022

Accurate 3D Body Shape Regression Using Metric and Semantic Attributes

CVPR 2022oral

While methods that regress 3D human meshes from images have progressed rapidly, the estimated body shapes often do not capture the true human shape. This is problematic since, for many applications, accurate body shape is as important as pose. The key reason that body shape accuracy lags pose accura…

Cited by 72PDFcodeScholar
2022

GOAL: Generating 4D Whole-Body Motion for Hand-Object Grasping

CVPR 2022poster

Generating digital humans that move realistically has many applications and is widely studied, but existing methods focus on the major limbs of the body, ignoring the hands and head. Hands have been separately studied, but the focus has been on generating realistic static grasps of objects. To synth…

Cited by 124PDFcodeScholar
2022

Human-Aware Object Placement for Visual Environment Reconstruction

CVPR 2022poster

Humans are in constant contact with the world as they move through it and interact with it. This contact is a vital source of information for understanding 3D humans, 3D scenes, and the interactions between them. In fact, we demonstrate that these human-scene interactions (HSIs) can be leveraged to…

Cited by 70PDFcodeScholar
2022

ICON: Implicit Clothed Humans Obtained From Normals

CVPR 2022poster

Current methods for learning realistic and animatable 3D clothed avatars need either posed 3D scans or 2D images with carefully controlled user poses. In contrast, our goal is to learn the avatar from only 2D images of people in unconstrained poses. Given a set of images, our method estimates a deta…

Cited by 343PDFcodeScholar
2022

SUPR: A Sparse Unified Part-Based Human Representation

ECCV 2022poster

"Statistical 3D shape models of the head, hands, and full body are widely used in computer vision and graphics. Despite their wide use, we show that existing models of the head and hands fail to capture the full range of motion for these parts. Moreover, existing work largely ignores the feet, which…

2021

Populating 3D Scenes by Learning Human-Scene Interaction

CVPR 2021poster

Humans live within a 3D space and constantly interact with it to perform tasks. Such interactions involve physical contact between surfaces that is semantically meaningful. Our goal is to learn how humans interact with scenes and leverage this to enable virtual characters to do the same. To that end…

Cited by 167PDFcodeScholar
2020

GRAB: A Dataset of Whole-Body Human Grasping of Objects

ECCV 2020poster

Training computers to understand, model, and synthesize human grasping requires a rich dataset containing complex 3D object shapes, detailed contact information, hand pose and shape, and the 3D body motion over time. While ""grasping"" is commonly thought of as a single hand stably lifting an object…

2020

Monocular Expressive Body Regression through Body-Driven Attention

ECCV 2020poster

To understand how people look, interact, or perform tasks, we need to quickly and accurately capture their 3D body, face, and hands together from an RGB image. Most existing methods focus only on parts of the body. A few recent approaches reconstruct full expressive 3D humans from images using 3D bo…

2019

Expressive Body Capture: 3D Hands, Face, and Body From a Single Image

CVPR 2019oral

To facilitate the analysis of human actions, interactions and emotions, we compute a 3D model of human body pose, hand pose, and facial expression from a single monocular image. To achieve this, we use thousands of 3D scans to train a new, unified, 3D model of the human body, SMPL-X, that extends SM…

Cited by 2080PDFcodeScholar
2019

Learning Joint Reconstruction of Hands and Manipulated Objects

CVPR 2019poster

Estimating hand-object manipulations is essential for in- terpreting and imitating human actions. Previous work has made significant progress towards reconstruction of hand poses and object shapes in isolation. Yet, reconstructing hands and objects during manipulation is a more challeng- ing task du…

Cited by 631PDFScholar
2019

Resolving 3D Human Pose Ambiguities With 3D Scene Constraints

ICCV 2019poster

To understand and analyze human behavior, we need to capture humans moving in, and interacting with, the world. Most existing methods perform 3D human pose estimation without explicitly considering the scene. We observe however that the world constrains the body and vice-versa. To motivate this, we…

Cited by 359PDFcodeScholar