← Search

Dechen Gao

3 accepted papers

2026

Look, Focus, Act: Efficient and Robust Robot Learning Via Human Gaze and Foveated Vision Transformers

ICRA 2026poster

Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on passive, uniform processing of raw camera images. In this work, we explore how in…

2026

VITA: Vision-to-Action Flow Matching Policy

ICLR 2026poster

Conventional flow matching and diffusion-based policies sample through iterative denoising from standard noise distributions (e.g., Gaussian), and require conditioning modules to repeatedly incorporate visual information during the generative process, incurring substantial time and memory overhead.…

Cited by 0SourcecodeScholar
2025

Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation

ICRA 2025

Imitation learning has demonstrated significant potential in performing high-precision manipulation tasks using visual feedback. However, it is common practice in imitation learning for cameras to be fixed in place, resulting in issues like occlusion and limited field of view. Furthermore, cameras a

Cited by 33SourcecodeScholar