NAACL 2025system demonstrations5 citations

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Jiatong Shi, Hye-jin Shim, Jinchuan Tian, Siddhant Arora, Haibin Wu, Darius Petermann, Jia Qi Yip, You Zhang

Abstract

In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA offers 65 metrics with 729 metric variations based on different configurations. These metrics encompass evaluations utilizing diverse external resources, including matching and non-matching reference audio, text transcriptions, and text captions. As a lightweight yet comprehensive toolkit, VERSA is versatile to support the evaluation of a wide range of downstream scenarios. To demonstrate its capabilities, this work highlights example use cases for VERSA, including audio coding, speech synthesis, speech enhancement, singing synthesis, and music generation. The toolkit is available at https://github.com/shinjiwlab/versa.

BibTeX
@inproceedings{shi-etal-2025-versa,
    title = "{VERSA}: A Versatile Evaluation Toolkit for Speech, Audio, and Music",
    author = "Shi, Jiatong  and
      Shim, Hye-jin  and
      Tian, Jinchuan  and
      Arora, Siddhant  and
      Wu, Haibin  and
      Petermann, Darius  and
      Yip, Jia Qi  and
      Zhang, You  and
      Tang, Yuxun  and
      Zhang, Wangyou  and
      Alharthi, Dareen Safar  and
      Huang, Yichen  and
      Saito, Koichi  and
      Han, Jionghao  and
      Zhao, Yiwen  and
      Donahue, Chris  and
      Watanabe, Shinji",
    editor = "Dziri, Nouha  and
      Ren, Sean (Xiang)  and
      Diao, Shizhe",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-demo.19/",
    pages = "191--209",
    ISBN = "979-8-89176-191-9"
}
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music · NAACL 2025