2025
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
ICRA 2025
Multimodal Large Language Models (MLLMs) have made significant progress in tasks such as image captioning and question answering. However, while these models can generate realistic captions, they often struggle with providing precise instructions, particularly when it comes to localizing and disambi