2025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
ICCV 2025poster
Multimodal large language models (MLLMs) excel at 2D visual understanding but remain limited in their ability to reason about 3D space. In this work, we leverage large-scale high-quality 3D scene data with open-set annotations to introduce 1) a novel supervised fine-tuning dataset and 2) a new evalu…