PicoAudio: Enabling Precise Temporal Controllability in Text-to-Audio Generation
Zeyu Xie, Xuenan Xu, Zhizheng Wu, Mengyue Wu
Abstract
Recently, audio generation tasks have attracted considerable research interests. Despite rapid advancements in generating high-fidelity audio that is coarsely aligned with the text description, precise temporal controllability is still a challenge, which is essential to integrate audio generation with real applications. In this work, we propose a temporal controlled audio generation framework, PicoAudio. It leverages data crawling, segmentation and filtering to simulate fine-grained temporally-aligned audio-text data. Furthermore, PicoAudio integrates temporal information to guide audio generation through tailored model design. With the effective text processing capabilities from large language models, PicoAudio can take natural language input and generate audio that aligns well with the temporal description in the input. Both subjective and objective evaluation demonstrate that PicoAudio dramatically surpasses current state-of-the-art generation models in terms of timestamp and occurrence frequency controllability. Generation samples are available at the $PicoAudio - Demo$.
BibTeX
@inproceedings{icassp2025_picoaudioenablin,
title = {PicoAudio: Enabling Precise Temporal Controllability in Text-to-Audio Generation},
author = {Zeyu Xie and Xuenan Xu and Zhizheng Wu and Mengyue Wu},
booktitle = {ICASSP 2025},
year = {2025}
}