← Search

Iman Soltani

6 accepted papers

2026

Look, Focus, Act: Efficient and Robust Robot Learning Via Human Gaze and Foveated Vision Transformers

ICRA 2026poster

Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on passive, uniform processing of raw camera images. In this work, we explore how in…

2026

RAMAC: Multimodal Risk-Aware Offline Reinforcement Learning and the Role of Behavior Regularization

ICML 2026poster

In safety-critical domains where online data collection is infeasible, offline reinforcement learning (RL) is attractive only if policies achieve high returns without catastrophic lower-tail risk. Prior work on risk-averse offline RL achieves safety at the cost of value- or model-based pessimism, an…

Cited by 0SourceScholar
2026

VITA: Vision-to-Action Flow Matching Policy

ICLR 2026poster

Conventional flow matching and diffusion-based policies sample through iterative denoising from standard noise distributions (e.g., Gaussian), and require conditioning modules to repeatedly incorporate visual information during the generative process, incurring substantial time and memory overhead.…

Cited by 0SourcecodeScholar
2025

Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation

ICRA 2025

Imitation learning has demonstrated significant potential in performing high-precision manipulation tasks using visual feedback. However, it is common practice in imitation learning for cameras to be fixed in place, resulting in issues like occlusion and limited field of view. Furthermore, cameras a

Cited by 33SourcecodeScholar
2024

Hierarchical End-to-End Autonomous Navigation Through Few-Shot Waypoint Detection

RA-L 2024

Human navigation is facilitated through the association of actions with landmarks, tapping into our ability to recognize salient features in our environment. Consequently, navigational instructions for humans can be extremely concise, such as short verbal descriptions, indicating a small memory requ

Cited by 10SourceScholar
2024

InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation

CoRL 2024poster

We present InterACT: Inter-dependency aware Action Chunking with Hierarchical Attention Transformers, a novel imitation learning framework for bimanual manipulation that integrates hierarchical attention to capture inter-dependencies between dual-arm joint states and visual inputs. InterACT consists…

Cited by 7SourceScholar