IJCAI 20260 citations

RIVS: Mitigating Hallucination in Large Vision-Language Models via Representation Intervention on Visual Grounding Shift

Xuanyu Yin, Xiaoye Qu, Wei Wei

Abstract

Large Vision-Language Models (LVLMs) demonstrate powerful generative capabilities yet remain prone to object hallucinations. Most existing methods mitigate this issue through training or decoding strategies, but provide limited exploration of how hallucinations arise from internal representations during generation. In this work, we study hallucination from the perspective of dynamic representation shift during generation and propose Representation Intervention based on Visual Grounding Shift (RIVS). Specifically, we first design an adaptive threshold-based method to identify visual attention drop points within the generation, finding that hallucinated tokens usually occur with decay of visual attention. Based on this observation, we construct a hallucination-related subspace from representation differences around these points on a small calibration set, without constructing contrast samples or extra supervision. During inference, we leverage the resulting hallucination-related subspace to perform an online projection-based intervention on intermediate hidden states to suppress the hallucination-related directions, mitigating hallucinations while preserving language quality. Our RIVS is training-free and computationally efficient. Experiments on four hallucination and two reasoning datasets demonstrate that RIVS consistently reduces hallucinations in both long and short sequence generation tasks.

AI Ethics, Trust, Fairnes: BiasAI Ethics, Trust, Fairnes: Trustworthy AIAI Ethics, Trust, Fairnes: Other
BibTeX
@inproceedings{ijcai2026_rivsmitigatingha,
  title = {RIVS: Mitigating Hallucination in Large Vision-Language Models via Representation Intervention on Visual Grounding Shift},
  author = {Xuanyu Yin and Xiaoye Qu and Wei Wei},
  booktitle = {IJCAI 2026},
  year = {2026}
}
RIVS: Mitigating Hallucination in Large Vision-Language Models via Representation Intervention on Visual Grounding Shift · IJCAI 2026