← Search

Maoyuan Shao

2 accepted papers

2026

CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language Misalignment

CVPR 2026

Vision-language models like CLIP have achieved remarkable progress in cross-modal representation learning, yet suffer from systematic misclassifications among visually and semantically similar categories. We observe that such confusion patterns are not random but persistently occur between specific

Cited by 0SourcecodeScholar
2025

Spotlighter: Revisiting Prompt Tuning from a Representative Mining View

EMNLP 2025

CLIP’s success has demonstrated that prompt tuning can achieve robust cross-modal semantic alignment for tasks ranging from open-domain recognition to fine-grained classification. However, redundant or weakly relevant feature components introduce noise and incur unnecessary computational costs. In t

Cited by 0SourcePDFScholar