2026
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
ICLR 2026poster
Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos. However, real-world embodied AI agents—such as robots and self…