AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
Zeyu Xie, Xuenan Xu, Zhizheng Wu, Mengyue Wu
Abstract
Recent advances in audio generation have enabled the creation of high-fidelity audio clips from free-form textual descriptions. However, temporal relation, a critical feature for audio content, is currently underrepresented in mainstream models, resulting in an imprecise temporal controllability. Specifically, users cannot accurately control the timestamps of sound events using free-form text. One significant challenge is the absence of a high-quality, temporally-aligned audio-text dataset, which is essential for training models with temporal control. The more temporally-aligned the annotations, the better the models can understand the precise relationship between audio outputs and temporal textual prompts. Therefore, we propose a temporally-aligned audio-text dataset, AudioTime. It provides text annotations rich in temporal information such as timestamps, duration, frequency, and ordering, covering almost all aspects of temporal control. Additionally, we offer a comprehensive test set and evaluation metric to assess the temporal control performance of text-to-audio generation models. Examples are available on the $AudioTime - Demo$.
BibTeX
@inproceedings{icassp2025_audiotimeatempor,
title = {AudioTime: A Temporally-aligned Audio-text Benchmark Dataset},
author = {Zeyu Xie and Xuenan Xu and Zhizheng Wu and Mengyue Wu},
booktitle = {ICASSP 2025},
year = {2025}
}