2025
Caption This, Reason That: VLMs Caught in the Middle
NeurIPS 2025spotlight
Vision-Language Models (VLMs) have shown remarkable progress in visual understanding in recent years. Yet, they still lag behind human capabilities in specific visual tasks such as counting or relational reasoning. To understand the underlying limitations, we adopt methodologies from cognitive scien…