2026
CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Vision-Language Model
ICLR 2026poster
Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descriptions, typically formed by inserting a class label (e.g., "cat") into a prompt (e.g., "a photo of a").Existing studies have shown that the score betwe…