2026
GeoMind: Explicit Spatial Reasoning via Dual-Reference Geometric Modeling
IJCAI 2026
While Vision-Language Models (VLMs) excel at semantic understanding, they struggle to comprehend 3D spatial relationships from limited views. Their reliance on implicit geometric encoding often leads to severe hallucinations and inconsistencies in spatial reasoning tasks. To address this, we introdu