2026
CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language Misalignment
CVPR 2026
Vision-language models like CLIP have achieved remarkable progress in cross-modal representation learning, yet suffer from systematic misclassifications among visually and semantically similar categories. We observe that such confusion patterns are not random but persistently occur between specific