← Search

Sombit Dey

4 accepted papers

2026

AR-VLA: Autoregressive Action Expert for Vision–Language–Action Models

RSS 2026poster

We propose a standalone autoregressive (AR) Action Expert that generates actions as a continuous causal sequence while conditioning on refreshable vision-language prefixes. In contrast to existing Vision-Language-Action (VLA) models and diffusion policies that reset temporal context with each new ob…

Cited by 0SourceScholar
2026

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding

CVPR 2026

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control.Yet their ability to generalize across new environments, tasks, and embodiments remains limited.We argue that a major bottleneck lies in their foundations: most RFMs are built by fine-tuning inter

Cited by 0SourceScholar
2025

ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models

ICRA 2025

Recent progress in large language models and access to large-scale robotic datasets has sparked a paradigm shift in robotics models transforming them into generalists able to adapt to various tasks, scenes, and robot modalities. A large step for the community are open Vision Language Action models w

Cited by 22SourceScholar
2023

Learning Whom to Trust in Navigation: Dynamically Switching Between Classical and Neural Planning

IROS 2023poster

Navigation of terrestrial robots is typically addressed either with localization and mapping (SLAM) followed by classical planning on the dynamically created maps, or by machine learning (ML), often through end-to-end training with reinforcement learning (RL) or imitation learning (IL). Recently, mo…

Cited by 5SourceScholar