2024
C-CLAPA: Improving Text-Audio Cross Domain Retrieval with Captioning and Augmentations
ICASSP 2024accepted
In this paper, we introduce Captioning decoder Contrastive Language-Audio Pretraining with data Augmantation (C-CLAPA), a new Audio-Text model for the Cross Domain Retrieval (CDR) task. The model’s backbone is comprised of two encoders, one for the text and the other for the audio. The embedding vec…