2025
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
ICASSP 2025accepted
We propose the joint speech translation and recognition (JSTAR) model that leverages the fast-slow cascaded encoder architecture for simultaneous end-to-end automatic speech recognition (ASR) and speech translation (ST). The model is transducer-based and uses a multi-objective training strategy that…