2025
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
ACL 2025finding
The recent advancement of Multimodal Large Language Models (MLLMs) has significantly improved their fine-grained perception of single images and general comprehension across multiple images. However, existing MLLMs still face challenges in achieving precise grounding in complex multi-image scenarios…