← Search

Jianfei Zhao

2 accepted papers

2026

Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention

CVPR 2026

Visual attention serves as the primary mechanism through which MLLMs interpret visual information; however, its limited localization capability often leads to hallucinations. We observe that although MLLMs can accurately extract visual semantics from visual tokens, they fail to fully leverage this a

Cited by 0SourceScholar
2025

Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow

EMNLP 2025

Due to the unidirectional masking mechanism, Decoder-Only models propagate information from left to right. LVLMs (Large Vision-Language Models) follow the same architecture, with visual information gradually integrated into semantic representations during forward propagation. Through systematic anal

Cited by 0SourcePDFScholar