2025
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
ACL 2025long
Vision-language models (VLMs) work well in tasks ranging from image captioning to visual question answering (VQA), yet they struggle with spatial reasoning, a key skill for understanding our physical world that humans excel at. We find that spatial relations are generally rare in widely used VL data…