← Search

Ruicong Liu

8 accepted papers

2026

DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation

CVPR 2026

Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without considering social context or modeling the mutual dynamics between

Cited by 0SourceScholar
2026

UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking

CVPR 2026

Generating lifelike conversational avatars requires modeling not just isolated speakers, but the dynamic, reciprocal interaction of speaking and listening.However, modeling the listener is exceptionally challenging: direct audio-driven training fails, producing stiff, static listening motions. This

Cited by 0SourcecodeScholar
2025

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance

ICCV 2025poster

This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individual within a 3D point cloud. Human inertial localization is challenging due to IM…

Cited by 0SourcePDFScholar
2024

ActionVOS: Actions as Prompts for Video Object Segmentation

ECCV 2024oral

"Delving into the realm of egocentric vision, the advancement of referring video object segmentation (RVOS) stands as pivotal in understanding human activities. However, existing RVOS task primarily relies on static attributes such as object names to segment target objects, posing challenges in dist…

2024

Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition

ECCV 2024poster

"Compared with visual signals, Inertial Measurement Units (IMUs) placed on human limbs can capture accurate motion signals while being robust to lighting variation and occlusion. While these characteristics are intuitively valuable to help egocentric action recognition, the potential of IMUs remains…

Cited by 9SourcePDFScholar
2024

Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation

CVPR 2024poster

The pursuit of accurate 3D hand pose estimation stands as a keystone for understanding human activity in the realm of egocentric vision. The majority of existing estimation methods still rely on single-view images as input leading to potential limitations e.g. limited field-of-view and ambiguity in…

2021

Generalizing Gaze Estimation With Outlier-Guided Collaborative Adaptation

ICCV 2021poster

Deep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons. In this paper, we propose a plug-and-play gaze adaptation fr…

Cited by 71PDFcodeScholar