← Search

Benhong Zhang

1 accepted papers

2026

GeoMind: Explicit Spatial Reasoning via Dual-Reference Geometric Modeling

IJCAI 2026

While Vision-Language Models (VLMs) excel at semantic understanding, they struggle to comprehend 3D spatial relationships from limited views. Their reliance on implicit geometric encoding often leads to severe hallucinations and inconsistencies in spatial reasoning tasks. To address this, we introdu

Cited by 0Scholar