2024
Discovering Unwritten Visual Classifiers with Large Language Models
ECCV 2024poster
"Multimodal pre-trained models, such as CLIP, are popular for zero-shot classification due to their open-vocabulary flexibility and high performance. However, vision-language models, which compute similarity scores between images and class labels, are largely black-box, with limited interpretability…