2026
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
ICLR 2026poster
Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet, most evaluations of VLMs focus on single-view settings, leaving their ability t…