← Search

Evonne Ng

7 accepted papers

2026

DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction

CVPR 2026

We present DuoMo, a generative method that recovers human motion in world-space coordinates from unconstrained videos with noisy or incomplete observations. Reconstructing such motion requires solving a fundamental trade-off: generalizing from diverse and noisy video inputs while maintaining global

Cited by 0SourcecodeScholar
2025

Pose Priors from Language Models

CVPR 2025poster

Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large multimodal models (LMMs) as priors for reconstructing contact poses, offering a scalable alternative to traditional meth…

2024

From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations

CVPR 2024poster

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio we output multiple possibilities of gestural motion for an individual including face body and hands. The key behind our method is…

2023

Can Language Models Learn to Listen?

ICCV 2023poster

We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the speaker's words with their timestamps, our approach autoregressively predicts a response of a listener: a sequence of lis…

Cited by 24PDFScholar
2022

Learning To Listen: Modeling Non-Deterministic Dyadic Facial Motion

CVPR 2022poster

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and speech audio of the speaker using a motion-audio cross attentio…

Cited by 105PDFScholar
2021

Body2Hands: Learning To Infer 3D Hands From Conversational Gesture Body Dynamics

CVPR 2021poster

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in non-verbal communication settings. We formulate the learning of th…

Cited by 55PDFcodeScholar
2020

You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions

CVPR 2020oral

The body pose of a person wearing a camera is of great interest for applications in augmented reality, healthcare, and robotics, yet much of the person's body is out of view for a typical wearable camera. We propose a learning-based approach to estimate the camera wearer's 3D body pose from egocentr…

Cited by 113PDFcodeScholar