2026
WAVLINK: COMPACT AUDIO–TEXT EMBEDDINGS WITH A GLOBAL WHISPER TOKEN
ICASSP 2026oral
Whisper has become the de-facto encoder for extracting general-purpose audio features in large audio-language models, where a 30-second clip is typically represented by 1500 frame features projected into an LLM. In contrast, audio-text embedding models like CLAP-based models have largely relied on a…