Investigating End-to-End ASR Architectures for Long Form Audio Transcription
Nithin Rao Koluguri, Samuel Kriman, Georgy Zelenfroind, Somshubra Majumdar, Dima Rekesh, Vahid Noroozi, Jagadeesh Balam, Boris Ginsburg
Abstract
This paper presents an overview and evaluation of some of the end-to-end ASR models on long-form audio. We study three categories of Automatic Speech Recognition(ASR) models based on their core architecture: (1) convolutional, (2) convolutional with squeeze-and-excitation, and (3) convolutional models with attention. We selected one ASR model from each category and evaluated the Word Error Rate, maximum audio length and real-time factor for each model on a variety of long audio benchmarks: Earnings-21 and 22, CORAAL, and TED-LIUM3. The model from the category of self-attention with local attention and global token has the best accuracy compared to other architectures. We also compared models with CTC and RNNT decoders and showed that CTC-based models are more robust and efficient than RNNT on long form audio.
BibTeX
@inproceedings{icassp2024_investigatingend,
title = {Investigating End-to-End ASR Architectures for Long Form Audio Transcription},
author = {Nithin Rao Koluguri and Samuel Kriman and Georgy Zelenfroind and Somshubra Majumdar and Dima Rekesh and Vahid Noroozi and Jagadeesh Balam and Boris Ginsburg},
booktitle = {ICASSP 2024},
year = {2024}
}