← Search

Zhida Qin

1 accepted papers

2025

From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection

ICML 2025poster

Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping, in alignment with fine-grained class descriptions generated by large language m…