2025
On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
ICASSP 2025accepted
Recent advances in audio-text cross-modal contrastive learning have shown its potential towards zero-shot learning. One possibility for this is by projecting item embeddings from pre-trained backbone neural networks into a cross-modal space in which item similarity can be calculated in either domain…