Mitigating Hallucinations in Multi-modal Large Language Models via Image Token Attention-Guided Decoding
Multi-modal large language models (MLLMs) integrate the inherent text generation capabilities of large language models with an understanding of other modalities, promising wide applications in open-ended tasks. Despite their success, they often generate plausible but incorrect content. This phenomen…