HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-fo