← Search

Rishabh Dabral

17 accepted papers

2026

EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents

CVPR 2026

Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting.However, existing capture systems typically rely on costly studio setups and wearable devices, limiting the large-scale c

Cited by 0SourcecodeScholar
2026

FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

CVPR 2026

We present FunREC, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunREC operates directly on in-t

Cited by 0SourcecodeScholar
2026

MIBURI: Towards Expressive Interactive Gesture Synthesis

CVPR 2026

Embodied Conversational Agents (ECAs) aim to emulate human face-to-face interaction through speech, gestures, and facial expressions. Current large language model (LLM)-based conversational agents lack embodiment and the expressive gestures essential for natural interaction. Existing solutions for E

Cited by 0SourcecodeScholar
2026

Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views

CVPR 2026

We present _Relightable Holoported Characters_ (RHC), a novel person-specific method for free-view rendering and relighting of full-body and highly dynamic humans solely observed from sparse-view RGB videos at inference. In contrast to classical one-light-at-a-time (OLAT)-based human relighting, our

Cited by 0SourceScholar
2026

SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokens

CVPR 2026

Synthesizing text-driven 3D human motion within realistic scenes requires learning both semantic intent ("walk to the couch") and physical feasibility (e.g., avoiding collisions). Current methods use generative frameworks that simultaneously learn high-level planning and low-level contact reasoning,

Cited by 0SourcecodeScholar
2025

Attention (as Discrete-Time Markov) Chains

NeurIPS 2025poster

We introduce a new interpretation of the attention matrix as a discrete-time Markov chain. Our interpretation sheds light on common operations involving attention scores such as selection, summation, and averaging in a unified framework. It further extends them by considering indirect attention, pro…

Cited by 0SourceScholar
2025

BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects

CVPR 2025poster

We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-base…

Cited by 2SourcePDFScholar
2025

Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input

CVPR 2025poster

This work focuses on tracking and understanding human motion using consumer wearable devices, such as VR/AR headsets, smart glasses, cellphones, and smartwatches. These devices provide diverse, multi-modal sensor inputs, including egocentric images, and 1-3 sparse IMU sensors in varied combinations.…

Cited by 0SourcePDFScholar
2025

FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video

CVPR 2025highlight

Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurat…

Cited by 0SourcePDFScholar
2025

Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures

CVPR 2025highlight

Real-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space to learn rendering primitives. However, they either jointly learn geometry and…

Cited by 2SourcePDFScholar
2025

Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis

CVPR 2025poster

Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically me…

Cited by 1SourcePDFScholar
2024

ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis

CVPR 2024poster

Gestures play a key role in human communication. Recent methods for co-speech gesture generation while managing to generate beat-aligned motions struggle generating gestures that are semantically aligned with the utterance. Compared to beat gestures that align naturally to the audio signal semantica…

Cited by 12SourcePDFScholar
2024

MetaCap: Meta-learning Priors from Multi-View Imagery for Sparse-view Human Performance Capture and Rendering

ECCV 2024poster

"Faithful human performance capture and free-view rendering from sparse RGB observations is a long-standing problem in Vision and Graphics. The main challenges are the lack of observations and the inherent ambiguities of the setting, e.g. occlusions and depth ambiguity. As a result, radiance fields,…

Cited by 9SourcePDFScholar
2024

ReMoS: 3D Motion-Conditioned Reaction Synthesis for Two-Person Interactions

ECCV 2024poster

"Current approaches for 3D human motion synthesis generate high-quality animations of digital humans performing a wide variety of actions and gestures. However, a notable technological gap exists in addressing the complex dynamics of multi-human interactions within this paradigm. In this work, we pr…

2023

Mofusion: A Framework for Denoising-Diffusion-Based Motion Synthesis

CVPR 2023highlight

Conventional methods for human motion synthesis have either been deterministic or have had to struggle with the trade-off between motion diversity vs motion quality. In response to these limitations, we introduce MoFusion, i.e., a new denoising-diffusion-based framework for high-quality conditional…

Cited by 187SourcePDFScholar
2021

Gravity-Aware Monocular 3D Human-Object Reconstruction

ICCV 2021poster

This paper proposes GraviCap, i.e., a new approach for joint markerless 3D human motion capture and object trajectory estimation from monocular RGB videos. We focus on scenes with objects partially observed during a free flight. In contrast to existing monocular methods, we can recover scale, object…

Cited by 31PDFScholar
2018

Learning 3D Human Pose from Structure and Motion

ECCV 2018poster

3D human pose estimation from a single image is a challenging problem, especially for in-the-wild settings due to the lack of 3D annotated data. We propose two anatomically inspired loss functions and use them with a weakly-supervised learning framework to jointly learn from large-scale in-the-wild…

Cited by 264SourcePDFScholar