2024
Improving Zero-Shot Generalization for CLIP with Variational Adapter
ECCV 2024poster
"The excellent generalization capability of pre-trained Vision-Language Models (VLMs) makes fine-tuning VLMs for downstream zero-shot tasks a popular choice. Despite achieving promising performance in the professionality of base classes, most existing fine-tuned methods suffer from feature confusion…