2022
Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword Units
ICASSP 2022accepted
In end-to-end automatic speech recognition (ASR), a model is expected to implicitly learn representations suitable for recognizing a word-level sequence. However, the huge abstraction gap between input acoustic signals and output linguistic tokens makes it challenging for a model to learn the repres…