AAAI 2026technical0 citations

Enhancing Spatial Reasoning Through Visual and Textual Thinking

Xun Liang, Xin Guo, Zhongming Jin, Weihang Pan, Penghui Shang, Deng Cai, Binbin Lin, Jieping Ye

Abstract

The spatial reasoning task aims to reason about the spatial relationships in 2D and 3D space, which is a fundamental capability for Visual Question Answering (VQA) and robotics. Although vision language models (VLMs) have developed rapidly in recent years, they are still struggling with the spatial reasoning task. In this paper, we introduce a method that can enhance Spatial reasoning through Visual and Textual thinking Simultaneously (SpatialVTS). In the spatial visual thinking phase, our model is trained to generate location-related specific tokens of important targets automatically. Not only are the objects mentioned in the problem addressed, but also the potential objects related to the reasoning are considered. During the spatial textual thinking phase, our model conducts long-term thinking based on visual cues and dialogues and gradually inferences the answers to spatial reasoning problems. To effectively support the model

BibTeX
@inproceedings{aaai2026_enhancingspatial,
  title = {Enhancing Spatial Reasoning Through Visual and Textual Thinking},
  author = {Xun Liang and Xin Guo and Zhongming Jin and Weihang Pan and Penghui Shang and Deng Cai and Binbin Lin and Jieping Ye},
  booktitle = {AAAI 2026},
  year = {2026}
}
Enhancing Spatial Reasoning Through Visual and Textual Thinking · AAAI 2026