2026
On the Intrinsic Limits of Transformer Image Embeddings in Non-Solvable Spatial Reasoning
ICML 2026poster
Vision Transformers (ViTs) excel in semantic recognition but exhibit systematic failures in spatial reasoning tasks such as mental rotation. While often attributed to data scale, this work argues that the limitation arises from the intrinsic circuit complexity of the architecture. By formalizing spa…