ICLR 2026poster0 citations

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

ZhiYuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du, Jiongrui Yan, Shi Shubin, Chengbo Yuan, Huizhi Liang

Abstract

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet, most evaluations of VLMs focus on single-view settings, leaving their ability to integrate multi-view information largely underexplored. At the same time, multi-camera setups are increasingly standard in robotic platforms, as they provide complementary perspectives to mitigate occlusion and depth ambiguity. Whether VLMs can effectively leverage such multi-view inputs for robotic reasoning therefore remains an open question. To bridge this gap, we introduce MV-RoboBench, a benchmark specifically designed to evaluate the multi-view spatial reasoning capabilities of VLMs in robotic manipulation. MV-RoboBench consists of 1.7k manually curated QA items across eight subtasks, divided into two primary categories: spatial understanding and robotic execution. We evaluate a diverse set of existing VLMs, including both open-source and closed-source models, along with enhanced versions augmented by Chain-of-Thought (CoT)-inspired enhancements. The results show that state-of-the-art models remain far below human performance, underscoring the substantial challenges VLMs face in multi-view robotic perception. Additionally, our analysis uncovers two key findings: (i) spatial intelligence and robotic task reasoning are correlated in multi-view robotic scenarios; and (ii) strong performance on existing general-purpose single-view spatial understanding benchmarks does not reliably translate to success in the robotic spatial tasks assessed by our benchmark. We release MV-RoboBench as an open resource to foster progress in spatially grounded VLMs and VLAs, providing a foundation for advancing embodied multi-view intelligence in robotics.

spatial understandingbenchmarkmulti-viewvlmrobotics
BibTeX
@inproceedings{
feng2026seeing,
title={Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes},
author={ZhiYuan Feng and Zhaolu Kang and Qijie Wang and Zhiying Du and Jiongrui Yan and Shi Shubin and Chengbo Yuan and Huizhi Liang and Yu Deng and Qixiu Li and Rushuai Yang and Ruichuan An and Leqi Zheng and Weijie Wang and Shawn Chen and Sicheng Xu and Yaobo Liang and Jiaolong Yang and Baining Guo},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=jXDZJAfRZB}
}