ICASSP 2025accepted0 citations

Stable Audio Open

Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons

Abstract

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model’s performance is competitive with the state-of-the-art across various metrics. Notably, the reported FD<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">openl3</inf> results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.

BibTeX
@inproceedings{icassp2025_stableaudioopen,
  title = {Stable Audio Open},
  author = {Zach Evans and Julian D. Parker and CJ Carr and Zack Zukowski and Josiah Taylor and Jordi Pons},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Stable Audio Open · ICASSP 2025