ICASSP 2025accepted0 citations

Coarse-to-Fine Text-to-Music Latent Diffusion

Luca A. Lanzendörfer, Tongyu Lu, Nathanaël Perraudin, Dorien Herremans, Roger Wattenhofer

Abstract

We introduce DiscoDiff, a text-to-music generative model that utilizes two latent diffusion models to produce high-fidelity 44.1kHz music hierarchically. Our approach significantly enhances audio quality through a coarse-to-fine generation strategy, leveraging residual vector quantization from the Descript Audio Codec. We consolidate this coarse-to-fine design through an important observation that the audio latent representation can be split into a primary and secondary part, controlling music content and details accordingly. We validate the effectiveness of our approach and text-audio alignment through various objective metrics. Furthermore, we provide access to high-quality synthetic captions for the MTG-Jamendo and FMA datasets, as well as open-sourcing DiscoDiff’s codebase and model checkpoints.

BibTeX
@inproceedings{icassp2025_coarsetofinetext,
  title = {Coarse-to-Fine Text-to-Music Latent Diffusion},
  author = {Luca A. Lanzendörfer and Tongyu Lu and Nathanaël Perraudin and Dorien Herremans and Roger Wattenhofer},
  booktitle = {ICASSP 2025},
  year = {2025}
}