2024
Compound Text-Guided Prompt Tuning via Image-Adaptive Cues
AAAI 2024technical
Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable generalization capabilities to downstream tasks. However, existing prompt tuning based frameworks need to parallelize learnable textual inputs for all categories, suffering from massive GPU memory consumption when there is a lar…