ACL 2022long26 citations

Learning When to Translate for Streaming Speech

Qian Dong, Yaoming Zhu, Mingxuan Wang, Lei Li

Abstract

How to find proper moments to generate partial sentence translation given a streaming speech input? Existing approaches waiting-and-translating for a fixed duration often break the acoustic units in speech, since the boundaries between acoustic units in speech are not even. In this paper, we propose MoSST, a simple yet effective method for translating streaming speech content. Given a usually long speech sequence, we develop an efficient monotonic segmentation module inside an encoder-decoder model to accumulate acoustic information incrementally and detect proper speech unit boundaries for the input in speech translation task. Experiments on multiple translation directions of the MuST-C dataset show that outperforms existing methods and achieves the best trade-off between translation quality (BLEU) and latency. Our code is available at https://github.com/dqqcasia/mosst.

BibTeX
@inproceedings{dong-etal-2022-learning,
    title = "Learning When to Translate for Streaming Speech",
    author = "Dong, Qian  and
      Zhu, Yaoming  and
      Wang, Mingxuan  and
      Li, Lei",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.50/",
    doi = "10.18653/v1/2022.acl-long.50",
    pages = "680--694"
}
Learning When to Translate for Streaming Speech · ACL 2022