← Search

Guangzong Si

3 accepted papers

2025

ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large Language Models

CVPR 2025poster

Contrastive decoding strategies are widely used to mitigate object hallucinations in multimodal large language models (MLLMs). By reducing over-reliance on language priors, these strategies ensure that generated content remains closely grounded in visual inputs, producing contextually accurate outpu…

2025

Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference

CVPR 2025poster

Multimodal large language models (MLLMs) improve performance on vision-language tasks by integrating visual features from pre-trained vision encoders into large language models (LLMs). However, how MLLMs process and utilize visual information remains unclear. In this paper, a shift in the dominant f…

2025

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

NeurIPS 2025poster

Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs). These methods work by constructing contrastive samples to induce hallucinations and then suppressing them in the output distribution. However, this paper demonstrates that suc…

Cited by 0SourceScholar