← Search

Yide Shentu

5 accepted papers

2026

EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations

ICRA 2026poster

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate head and hand movements, continuously reposition their viewpoin…

2024

From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control

IROS 2024poster

Hierarchical control for robotics has long been plagued by the need to have a well defined interface layer to communicate between high-level task planners and low-level policies. With the advent of LLMs, language has been emerging as a prospective interface layer. However, this has several limitatio…

Cited by 11SourceScholar
2024

GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators

IROS 2024poster

Humans can teleoperate robots to accomplish complex manipulation tasks. Imitation learning has emerged as a powerful framework that leverages human teleoperated demonstrations to teach robots new skills. However, the performance of the learned policies is bottlenecked by the quality, scale, and vari…

Cited by 106SourcecodeScholar
2022

Autoregressive Uncertainty Modeling for 3D Bounding Box Prediction

ECCV 2022poster

"3D bounding boxes are a widespread intermediate representation in many computer vision applications. However, predicting them is a challenging task, largely due to partial observability, which motivates the need for a strong sense of uncertainty. While many recent methods have explored better archi…

Cited by 7SourcePDFScholar
2018

Zero-Shot Visual Imitation

ICLR 2018oral

The current dominant paradigm for imitation learning relies on strong supervision of expert actions to learn both 'what' and 'how' to imitate. We pursue an alternative paradigm wherein an agent first explores the world without any expert supervision and then distills its experience into a goal-condi…