ACL 2025finding0 citations

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Wenxi Chen, Ziyang Ma, Ruiqi Yan, Yuzhe Liang, Xiquan Li, Ruiyang Xu, Zhikang Niu, Yanqiao Zhu

Abstract

Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system with single-stage training. SLAM-Omni achieves zero-shot timbre control by modeling spoken language with semantic tokens and decoupling speaker information to a vocoder. By predicting grouped speech semantic tokens at each step, our method significantly reduces the sequence length of audio tokens, accelerating both training and inference. Additionally, we propose historical text prompting to compress dialogue history, facilitating efficient multi-round interactions. Comprehensive evaluations reveal that SLAM-Omni outperforms prior models of similar scale, requiring only 15 hours of training on 4 GPUs with limited data. Notably, it is the first spoken dialogue system to achieve competitive performance with a single-stage training approach, eliminating the need for pre-training on TTS or ASR tasks. Further experiments validate its multilingual and multi-turn dialogue capabilities on larger datasets.

BibTeX
@inproceedings{chen-etal-2025-slam,
    title = "{SLAM}-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training",
    author = "Chen, Wenxi  and
      Ma, Ziyang  and
      Yan, Ruiqi  and
      Liang, Yuzhe  and
      Li, Xiquan  and
      Xu, Ruiyang  and
      Niu, Zhikang  and
      Zhu, Yanqiao  and
      Yang, Yifan  and
      Liu, Zhanxun  and
      Yu, Kai  and
      Hu, Yuxuan  and
      Li, Jinyu  and
      Lu, Yan  and
      Liu, Shujie  and
      Chen, Xie",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.115/",
    doi = "10.18653/v1/2025.findings-acl.115",
    pages = "2262--2282",
    ISBN = "979-8-89176-256-5"
}
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training · ACL 2025