2025
From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
ICML 2025poster
Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping, in alignment with fine-grained class descriptions generated by large language m…