Mitigating Hallucination in Vision-Language Model with Depth and Spatial-aware Key-Value Refinement
Large vision–language models (VLMs) deliver state-of-the-art results on a wide range of multimodal tasks, yet they remain prone to visual hallucinations, producing content that is not grounded in the input image. Despite progress with visual supervision, reinforcement learning, and post-hoc attenti…