2026
SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contexts
CVPR 2026
Large-scale pre-trained image-text models exhibit robust multimodal representation, yet applying contrastive language-image pretraining (CLIP) to audio-visual localization remains challenging. Replacing the classification token ([CLS]) with an audio-embedded token ([V_A])struggles to capture semanti