← Search

Huatian Zhang

4 accepted papers

2026

Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models

CVPR 2026

Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) by learning from preference pairs. One of its key challenges lies in how to transfer the sequence-level preference into fine-grained supervision on vis

Cited by 0SourcecodeScholar
2024

Homology Consistency Constrained Efficient Tuning for Vision-Language Models

NeurIPS 2024poster

Efficient transfer learning has shown remarkable performance in tuning large-scale vision-language models (VLMs) toward downstream tasks with limited data resources. The key challenge of efficient transfer lies in adjusting image-text alignment to be task-specific while preserving pre-trained genera…

Cited by 0SourcePDFScholar
2024

Identification of Necessary Semantic Undertakers in the Causal View for Image-Text Matching

AAAI 2024technical

Image-text matching bridges vision and language, which is a fundamental task in multimodal intelligence. Its key challenge lies in how to capture visual-semantic relevance. Fine-grained semantic interactions come from fragment alignments between image regions and text words. However, not all fragmen…

2022

Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text Matching

AAAI 2022technical

Image-text matching bridges vision and language, which is a crucial task in the field of multi-modal intelligence. The key challenge lies in how to measure image-text relevance accurately as matching evidence. Most existing works aggregate the local semantic similarities of matched region-word pairs…