2024
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models
EMNLP 2024finding
Fine-grained image classification, especially in zero-/few-shot scenarios, poses a considerable challenge for vision-language models (VLMs) like CLIP, which often struggle to differentiate between semantically similar classes due to insufficient supervision for fine-grained tasks. On the other hand,…