IJCAI 20260 citations

DeepL Voice: Real-Time Speech-to-Speech Translation

Johannes Ernesti, Peter Kaiser, Jonas Heinze, Elnaz Shafaei-Bajestan, Kristina Geißler, Weiyue Wang, Johannes Beck, Sascha Brinker

Abstract

DeepL Voice is a real-time speech-to-speech translation system for global business communication, following a pragmatic incremental approach: developing a production-grade cascaded speech-to-speech-translation (S2ST) system, while exploring end-to-end solutions in parallel. The production system (launched November 2024) achieves competitive transcription quality through proprietary real-time ASR models and eliminates translation "flickering" via stable text streaming while maintaining low latency. Supporting 18 input languages and 30+ target languages, it offers DeepL Voice for Meetings (Microsoft Teams/Zoom integration) and DeepL Voice for Conversations (mobile apps). Key features include customizable formality and glossary support for business-appropriate communication, with voice cloning TTS under development. Demo Video: https://youtu.be/DMMcti2f4rc

BibTeX
@inproceedings{ijcai2026_deeplvoicerealti,
  title = {DeepL Voice: Real-Time Speech-to-Speech Translation},
  author = {Johannes Ernesti and Peter Kaiser and Jonas Heinze and Elnaz Shafaei-Bajestan and Kristina Geißler and Weiyue Wang and Johannes Beck and Sascha Brinker and Thorben Finke},
  booktitle = {IJCAI 2026},
  year = {2026}
}
DeepL Voice: Real-Time Speech-to-Speech Translation · IJCAI 2026