ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
Siddhant Arora, Yifan Peng, Jiatong Shi, Jinchuan Tian, William Chen, Shikhar Bharadwaj, Hayato Futami, Yosuke Kashiwagi
Abstract
Advancements in audio foundation models (FMs) have fueled interest in end-to-end (E2E) spoken dialogue systems, but different web interfaces for each system makes it challenging to compare and contrast them effectively. Motivated by this, we introduce an open-source, user-friendly toolkit designed to build unified web interfaces for various cascaded and E2E spoken dialogue systems. Our demo further provides users with the option to get on-the-fly automated evaluation metrics such as (1) latency, (2) ability to understand user input, (3) coherence, diversity, and relevance of system response, and (4) intelligibility and audio quality of system output. Using the evaluation metrics, we compare various cascaded and E2E spoken dialogue systems with a human-human conversation dataset as a proxy. Our analysis demonstrates that the toolkit allows researchers to effortlessly compare and contrast different technologies, providing valuable insights such as current E2E systems having poorer audio quality and less diverse responses. An example demo produced using our toolkit is publicly available here: https://huggingface.co/spaces/Siddhant/Voice_Assistant_Demo.
BibTeX
@inproceedings{arora-etal-2025-espnet,
title = "{ESP}net-{SDS}: Unified Toolkit and Demo for Spoken Dialogue Systems",
author = "Arora, Siddhant and
Peng, Yifan and
Shi, Jiatong and
Tian, Jinchuan and
Chen, William and
Bharadwaj, Shikhar and
Futami, Hayato and
Kashiwagi, Yosuke and
Tsunoo, Emiru and
Shimizu, Shuichiro and
Srivastav, Vaibhav and
Watanabe, Shinji",
editor = "Dziri, Nouha and
Ren, Sean (Xiang) and
Diao, Shizhe",
booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.naacl-demo.21/",
pages = "248--259",
ISBN = "979-8-89176-191-9"
}