← Search

Fei Ming

1 accepted papers

2026

IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination

ICML 2026poster

Hallucination remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While Direct Preference Optimization (DPO) is a key alignment framework, existing approaches often rely heavily on costly external evaluators for scoring or rewriting, incurring off-policy learnability gaps a…

Cited by 0SourceScholar