2025
iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs
EMNLP 2025
Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal benchmark designed to evaluate the spatial reasoning capabilities of VLMs acting as agents. iVISPAR is based on a varia