Phone-Informed Refinement of Synthesized Mel Spectrogram for Data Augmentation in Speech Recognition
While recent end-to-end automatic speech recognition (ASR) models achieve high performance, we need to prepare an abundant amount of training data, which is a barrier to apply them to a specific domain. To mitigate the lack of training data, text-to-speech (TTS) systems have been utilized to leverag…