2026
Seeing Without Understanding: Disentangling Perception, Reasoning, and Simulation in VLM Gameplay
ICML 2026poster
While Vision-Language Models (VLMs) excel on static visual benchmarks, they consistently underperform in game-based reasoning environments. Existing evaluations conflate failures in perception, rule comprehension, and reasoning. We propose a two-stage diagnostic framework that decomposes VLM perform…