← Search

Ian Chuang

4 accepted papers

2026

Look, Focus, Act: Efficient and Robust Robot Learning Via Human Gaze and Foveated Vision Transformers

ICRA 2026poster

Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on passive, uniform processing of raw camera images. In this work, we explore how in…

2026

VITA: Vision-to-Action Flow Matching Policy

ICLR 2026poster

Conventional flow matching and diffusion-based policies sample through iterative denoising from standard noise distributions (e.g., Gaussian), and require conditioning modules to repeatedly incorporate visual information during the generative process, incurring substantial time and memory overhead.…

Cited by 0SourcecodeScholar
2025

Mechanistic Interpretability for Steering Vision-Language-Action Models

CoRL 2025poster

Vision-Language-Action (VLA) models are a promising path to realizing generalist embodied agents that can quickly adapt to new tasks, modalities, and environments. However, methods for interpreting and steering VLAs fall far short of classical robotics pipelines, which are grounded in explicit model…

Cited by 0SourceScholar
2024

InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation

CoRL 2024poster

We present InterACT: Inter-dependency aware Action Chunking with Hierarchical Attention Transformers, a novel imitation learning framework for bimanual manipulation that integrates hierarchical attention to capture inter-dependencies between dual-arm joint states and visual inputs. InterACT consists…

Cited by 7SourceScholar