RSS 2025poster18 citations

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models

Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Jiayuan Gu, Zhigang Wang, Yan Ding

Abstract

In this paper, we claim that spatial understanding is the keypoint in robot manipulation, and propose SpatialVLA to explore effective spatial representations for the robot foundation model. Specifically, we propose Ego3D Position Encoding to inject 3D information into VLA’s input observations, and introduce

BibTeX
@inproceedings{rss2025_spatialvlaexplor,
  title = {SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models},
  author = {Delin Qu and Haoming Song and Qizhi Chen and Yuanqi Yao and Xinyi Ye and Jiayuan Gu and Zhigang Wang and Yan Ding and Bin Zhao and Dong Wang and Xuelong Li},
  booktitle = {RSS 2025},
  year = {2025}
}
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models · RSS 2025