← Search

Zhuoheng Li

3 accepted papers

2026

RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning

CVPR 2026

We introduce Robowheel, a data engine that converts human hand-object interaction (HOI) videos into training-ready supervision for cross-morphology robotic learning. From monocular RGB/RGB-D inputs, we perform high-precision HOI reconstruction and enforce physical plausibility via a reinforcement le

Cited by 0SourceScholar
2025

Acknowledging Focus Ambiguity in Visual Questions

ICCV 2025poster

No published work on visual question answering (VQA) accounts for ambiguity regarding where the content described in the question is located in the image. To fill this gap, we introduce VQ-FocusAmbiguity, the first VQA dataset that visually grounds each plausible image region a question could refer…

Cited by 0SourcePDFScholar
2025

HumanMM: Global Human Motion Recovery from Multi-shot Videos

CVPR 2025poster

In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to applications such as motion generation and motion understan…