2021
Construction of a Large-Scale Japanese ASR Corpus on TV Recordings
ICASSP 2021accepted
This paper presents a new large-scale Japanese speech corpus for training automatic speech recognition (ASR) systems. This corpus contains over 2,000 hours of speech with transcripts built on Japanese TV recordings and their subtitles. We develop herein an iterative workflow to extract matching audi…