← Search

Yuan-Ming Li

8 accepted papers

2026

Beyond Mimicry: Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations

CVPR 2026

Enabling humanoid robots to physically interact with humans is a critical frontier, but progress is hindered by the scarcity of high-quality Human-Humanoid Interaction (HHoI) data. While leveraging abundant Human-Human Interaction (HHI) data presents a scalable alternative, we first demonstrate that

Cited by 0SourceScholar
2026

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

CVPR 2026

We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual prompts, such as masks or points, SWIM leverages mask supervis

Cited by 0SourcecodeScholar
2025

Less Static, More Private: Towards Transferable Privacy-Preserving Action Recognition by Generative Decoupled Learning

ICCV 2025poster

This work focuses on the task of privacy-preserving action recognition (PPAR), which aims to protect individual privacy in action videos without compromising recognition performance. Despite recent advancements, existing PPAR models still struggle with video domain shifts. To address this challenge,…

Cited by 0SourcePDFScholar
2025

Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks

CVPR 2025poster

Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering errors or rely on static prototypes to represent normal actions. However, these approaches typically overlook the common sce…

2025

ViSpeak: Visual Instruction Feedback in Streaming Videos

ICCV 2025poster

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming…

2024

EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding

ECCV 2024poster

"We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-body action understanding datasets, EgoExo-Fitness not only contains videos from…