2026
GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
CVPR 2026
Visual grounding, localizing objects from natural language descriptions, represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on existing benchmarks, a fundamental question remains: can MLLMs truly visually g