2025
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
NeurIPS 2025spotlight
Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or 2.5D data to incorporate spatial awareness, restricting the…