← Search

Changhao He

4 accepted papers

2026

Bootstrapping Multi-view Learning for Test-time Noisy Correspondence

CVPR 2026

Multi-view learning fuses complementary views to improve perception, but real-world deployments often suffer from Test-time Noisy Correspondence (TNC) -- cross-view misalignment caused by asynchronous sampling, transient network congestion, or other disturbances. Such misalignment introduces semanti

Cited by 0SourcecodeScholar
2026

DOUBT: Decoupled Object-level Understanding and Bridging via vMF-based Trustworthiness for Hallucination Detection in MLLMs

ICML 2026oral

Multimodal Large Language Models (MLLMs) frequently produce hallucinations (i.e., assertions that contradict the image or facts), undermining reliability in high-risk applications. Existing detection approaches typically feed images and texts jointly and estimate hallucination scores by measuring th…

Cited by 0SourceScholar
2026

RLSF-V: Mitigating Hallucinations in MLLMs via Fuzzy Semantic Self-Feedback

ICML 2026poster

Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception for open-world understanding, but exacerbate LLMs' hallucinations, in which generated text contradicts visual evidence or common sense. To mitigate hallucinations, a dominant strategy is Direct Prefer…

Cited by 0SourceScholar
2025

Learning with Noisy Triplet Correspondence for Composed Image Retrieval

CVPR 2025poster

Composed Image Retrieval (CIR) enables editable image search by integrating a query pair--a reference image ref and a textual modification mod--to retrieve a target image tar that reflects the intended change. While existing CIR methods have shown promising performance using well-annotated triplets…