RSS 2025poster18 citations
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models
Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Jiayuan Gu, Zhigang Wang, Yan Ding
Abstract
In this paper, we claim that spatial understanding is the keypoint in robot manipulation, and propose SpatialVLA to explore effective spatial representations for the robot foundation model. Specifically, we propose Ego3D Position Encoding to inject 3D information into VLA’s input observations, and introduce
BibTeX
@inproceedings{rss2025_spatialvlaexplor,
title = {SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models},
author = {Delin Qu and Haoming Song and Qizhi Chen and Yuanqi Yao and Xinyi Ye and Jiayuan Gu and Zhigang Wang and Yan Ding and Bin Zhao and Dong Wang and Xuelong Li},
booktitle = {RSS 2025},
year = {2025}
}