Coarse-to-Fine Text-to-Music Latent Diffusion
Luca A. Lanzendörfer, Tongyu Lu, Nathanaël Perraudin, Dorien Herremans, Roger Wattenhofer
Abstract
We introduce DiscoDiff, a text-to-music generative model that utilizes two latent diffusion models to produce high-fidelity 44.1kHz music hierarchically. Our approach significantly enhances audio quality through a coarse-to-fine generation strategy, leveraging residual vector quantization from the Descript Audio Codec. We consolidate this coarse-to-fine design through an important observation that the audio latent representation can be split into a primary and secondary part, controlling music content and details accordingly. We validate the effectiveness of our approach and text-audio alignment through various objective metrics. Furthermore, we provide access to high-quality synthetic captions for the MTG-Jamendo and FMA datasets, as well as open-sourcing DiscoDiff’s codebase and model checkpoints.
BibTeX
@inproceedings{icassp2025_coarsetofinetext,
title = {Coarse-to-Fine Text-to-Music Latent Diffusion},
author = {Luca A. Lanzendörfer and Tongyu Lu and Nathanaël Perraudin and Dorien Herremans and Roger Wattenhofer},
booktitle = {ICASSP 2025},
year = {2025}
}