2026
P$^2$-DPO:Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization
ICLR 2026poster
Hallucination has recently garnered significant research attention in Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) aims to learn directly from the corrected preferences provided by humans, thereby addressing the hallucination issue. Despite its success, this paradigm ha…