← Search

Shuo Cheng

14 accepted papers

2025

EgoMimic: Scaling Imitation Learning via Egocentric Video

ICRA 2025

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through

Cited by 136SourcecodeScholar
2025

Generalizable Domain Adaptation for Sim-and-Real Policy Co-Training

NeurIPS 2025poster

Behavior cloning has shown promise for robot manipulation, but real-world demonstrations are costly to acquire at scale. While simulated data offers a scalable alternative, particularly with advances in automated demonstration generation, transferring policies to the real world is hampered by variou…

Cited by 0SourceScholar
2024

Coarse-to-Fine Detection of Multiple Seams for Robotic Welding

IROS 2024poster

Efficiently detecting target weld seams while ensuring sub-millimeter accuracy has always been an important challenge in autonomous welding, which has significant application in industrial practice. Previous works mostly focused on recognizing and localizing welding seams one by one, leading to infe…

Cited by 0SourceScholar
2024

NOD-TAMP: Generalizable Long-Horizon Planning with Neural Object Descriptors

CoRL 2024poster

Solving complex manipulation tasks in household and factory settings remains challenging due to long-horizon reasoning, fine-grained interactions, and broad object and scene diversity. Learning skills from demonstrations can be an effective strategy, but such methods often have limited generalizabil…

Cited by 0SourcecodeScholar
2023

Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning

CoRL 2023poster

Practical Imitation Learning (IL) systems rely on large human demonstration datasets for successful policy learning. However, challenges lie in maintaining the quality of collected data and addressing the suboptimal nature of some demonstrations, which can compromise the overall dataset quality and…

Cited by 9SourceScholar
2020

Deep Stereo Using Adaptive Thin Volume Representation With Uncertainty Awareness

CVPR 2020oral

We present Uncertainty-aware Cascaded Stereo Network (UCS-Net) for 3D reconstruction from multiple RGB images. Multi-view stereo (MVS) aims to reconstruct fine-grained scene geometry from multi-view images. Previous learning-based MVS methods estimate per-view depth using plane sweep volumes (PSVs)…

Cited by 383PDFScholar
2018

Fine-Grained Video Captioning for Sports Narrative

CVPR 2018poster

Despite recent emergence of video caption methods, how to generate fine-grained video descriptions (i.e., long and detailed commentary about individual movements of multiple subjects as well as their frequent interactions) is far from being solved, which however has great applications such as automa…

Cited by 76SourcePDFScholar
2018

Pose Transferrable Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is an important task in the field of intelligent security. A key challenge is how to capture human pose variations, while existing benchmarks (i.e., Market1501, DukeMTMC-reID, CUHK03, etc.) do NOT provide sufficient pose coverage to train a robust ReID system. To add…

Cited by 456SourcePDFScholar