2024
FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance
ECCV 2024poster
"CLIP has achieved impressive zero-shot performance after pretraining on a large-scale dataset consisting of paired image-text data. Previous works have utilized CLIP by incorporating manually designed visual prompts like colored circles and blur masks into the images to guide the model’s attention,…