2024
Fast Timing-Conditioned Latent Audio Diffusion
ICML 2024oral
Generating long-form 44.1kHz stereo audio from text prompts can be computationally demanding. Further, most previous works do not tackle that music and sound effects naturally vary in their duration. Our research focuses on the efficient generation of long-form, variable-length stereo music and soun…