2024
Mtdiffusion: Multi-Task Diffusion Model With Dual-Unet for Foley Sound Generation
ICASSP 2024accepted
It is a common method to quantify the latent in audio generation and then use diffusion models to estimate noise or data from the corrupted data to generate the quantized latent. Unlike the method, we consider that the targets estimated by the diffusion model include both noise and data, rather than…