ICASSP 2025accepted0 citations

Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions

Yi Yuan, Dongya Jia, Xiaobin Zhuang, Yuanzhe Chen, Zhuo Chen, Yuping Wang, Yuxuan Wang, Xubo Liu

Abstract

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work aims to create a large-scale audio dataset with rich captions for improving audio generation models. We first develop an automated pipeline to generate detailed captions by transforming predicted visual captions, audio captions, and tagging labels into comprehensive descriptions using a Large Language Model (LLM). The resulting dataset, Sound-VECaps, comprises 1.66M high-quality audio-caption pairs with enriched details including audio event orders, occurred places and environment information. We then demonstrate that training the text-to-audio generation models with Sound-VECaps significantly improves the performance on complex prompts. Furthermore, we conduct ablation studies of the models on several downstream audio-language tasks, showing the potential of Sound-VECaps in advancing audio-text representation learning.Dataset and demos are available at https://yyua8222.github.io/Sound-VECaps-demo/.

BibTeX
@inproceedings{icassp2025_soundvecapsimpro,
  title = {Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions},
  author = {Yi Yuan and Dongya Jia and Xiaobin Zhuang and Yuanzhe Chen and Zhuo Chen and Yuping Wang and Yuxuan Wang and Xubo Liu and Xiyuan Kang and Mark D. Plumbley and Wenwu Wang},
  booktitle = {ICASSP 2025},
  year = {2025}
}