SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
Large vision-language models such as CLIP have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models often develop multimodal spurious biases, the undesirable tendency to rely on spurious features. For example, CLIP may infer