← Search

Rongbo Luan

1 accepted papers

2025

Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment

NeurIPS 2025poster

Despite Contrastive Language–Image Pre-training (CLIP)'s remarkable capability to retrieve content across modalities, a substantial modality gap persists in its feature space. Intriguingly, we discover that off-the-shelf MLLMs (Multimodal Large Language Models) demonstrate powerful inherent modality…

Cited by 0SourceScholar