DAPE-BR: Distance-Aware Positional Encoding for Mitigating Object Hallucination in LVLMs
Mingrui Xie, Tianxiang Xu, Qianhai Tang, Shanming Yao, Xiaofeng Zhang, Junliang Du
Abstract
Large Vision–Language Models (LVLMs) have garnered substantial interest owing to their impressive ability to interpret visual inputs and converse with users.Nevertheless, LVLMs still suffer from object hallucination – generating descriptions for objects that are absent from the image, which undermines reliability and hinders real-world deployment. We propose DAPE-BR, a positional-alignment scheme that (i) preserves the pretrained weight order while globally—- visual–text distances, (ii) embeds an isotropic fused patch-distance metric, and (iii) applies a patch-distance causal mask to enforce spatial causality. Extensive experiments on POPE, MMStar and SQA show that DAPE-BR consistently reduces hallucinations and boosts.
BibTeX
@inproceedings{emnlp2025_dapebrdistanceaw,
title = {DAPE-BR: Distance-Aware Positional Encoding for Mitigating Object Hallucination in LVLMs},
author = {Mingrui Xie and Tianxiang Xu and Qianhai Tang and Shanming Yao and Xiaofeng Zhang and Junliang Du},
booktitle = {EMNLP 2025},
year = {2025}
}