← Search

Omid Taheri

9 accepted papers

2026

CLUTCH: Contextualized Language model for Unlocking Text-Conditioned Hand motion modelling in the wild

ICLR 2026poster

Hands play a central role in daily life, yet modeling natural hand motions remains underexplored. Existing methods that tackle text-to-hand-motion generation or hand animation captioning rely on studio-captured datasets with limited actions and contexts, making them costly to scale to “in-the-wild”…

Cited by 0SourceScholar
2025

InteractVLM: 3D Interaction Reasoning from 2D Foundational Models

CVPR 2025poster

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth ambiguities, and widely varying object shapes. Existing methods r…

2024

HUMOS: Human Motion Model Conditioned on Body Shape

ECCV 2024poster

"Generating realistic human motion is crucial for many computer vision and graphics applications. The rich diversity of human body shapes and sizes significantly influences how people move. However, existing motion models typically overlook these differences, using a normalized, average body instead…

2024

WANDR: Intention-guided Human Motion Generation

CVPR 2024poster

Synthesizing natural human motions that enable a 3D human avatar to walk and reach for arbitrary goals in 3D space remains an unsolved problem with many applications. Existing methods (data-driven or using reinforcement learning) are limited in terms of generalization and motion naturalness. A prima…

Cited by 12SourcePDFScholar
2023

3D Human Pose Estimation via Intuitive Physics

CVPR 2023poster

Estimating 3D humans from images often produces implausible bodies that lean, float, or penetrate the floor. Such methods ignore the fact that bodies are typically supported by the scene. A physics engine can be used to enforce physical plausibility, but these are not differentiable, rely on unreali…

Cited by 98SourcePDFScholar
2023

ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

CVPR 2023poster

Humans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because there exist no datasets with ground-truth 3D annotations for…

2022

GOAL: Generating 4D Whole-Body Motion for Hand-Object Grasping

CVPR 2022poster

Generating digital humans that move realistically has many applications and is widely studied, but existing methods focus on the major limbs of the body, ignoring the hands and head. Hands have been separately studied, but the focus has been on generating realistic static grasps of objects. To synth…

Cited by 124PDFcodeScholar
2020

GRAB: A Dataset of Whole-Body Human Grasping of Objects

ECCV 2020poster

Training computers to understand, model, and synthesize human grasping requires a rich dataset containing complex 3D object shapes, detailed contact information, hand pose and shape, and the 3D body motion over time. While ""grasping"" is commonly thought of as a single hand stably lifting an object…