DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignment of text, video, and audio embeddings via a single contrastive loss. However,…