FAM: Fine-Grained Alignment Matters in Multimodal Embedding Learning with Large Vision-Language Models
Learning multimodal representation is a fundamental task that supports a wide range of applications such as visual-text retrieval. While pioneering approaches e.g., CLIP paves the way by learning separated encoders for different modalities, they struggle to model complex interactions between modalit