2026
SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
ICML 2026poster
We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio seque…