← Search

Ikhsanul Habibie

5 accepted papers

2024

ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis

CVPR 2024poster

Gestures play a key role in human communication. Recent methods for co-speech gesture generation while managing to generate beat-aligned motions struggle generating gestures that are semantically aligned with the utterance. Compared to beat gestures that align naturally to the audio signal semantica…

Cited by 12SourcePDFScholar
2023

Imitator: Personalized Speech-driven 3D Facial Animation

ICCV 2023poster

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input audio without considering the identity-specific speaking st…

Cited by 60PDFcodeScholar
2021

Monocular Real-Time Full Body Capture With Inter-Part Correlations

CVPR 2021poster

We present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network architecture that exploits correlations between body and hands at high computational e…

Cited by 72PDFScholar
2020

Monocular Real-Time Hand Shape and Motion Capture Using Multi-Modal Data

CVPR 2020poster

We present a novel method for monocular hand shape and pose estimation at unprecedented runtime performance of 100fps and at state-of-the-art accuracy. This is enabled by a new learning based architecture designed such that it can make use of all the sources of available hand training data: image da…

Cited by 257PDFcodeScholar
2019

In the Wild Human Pose Estimation Using Explicit 2D Features and Intermediate 3D Representations

CVPR 2019oral

Convolutional Neural Network based approaches for monocular 3D human pose estimation usually require a large amount of training images with 3D pose annotations. While it is feasible to provide 2D joint annotations for large corpora of in-the-wild images with humans, providing accurate 3D annotations…

Cited by 178PDFScholar