2025
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
NeurIPS 2025poster
Despite recent advances, vision-language models trained with standard contrastive objectives still struggle with compositional reasoning -- the ability to understand structured relationships between visual and linguistic elements. This shortcoming is largely due to the tendency of the text encoder t…