Speech analysis of sung-speech and lyric recognition in monophonic singing
Dairoku Kawai, Kazumasa Yamamoto, Seiichi Nakagawa
Abstract
Lyric recognition in singing is challenging because of a number of problems, including a lack of singing databases, superposed musical instruments and different spectral variations. First of all, we investigated the difference of spectral variations among read speech, spontaneous speech and sung speech and we found that sung speech recognition was the most difficult. Next, we consider Japanese lyric recognition in monophonic singing that contains no musical instruments. To express singing well, we use an n-gram language model with a lyrics corpus, singing-adapted acoustic models, and plural pronunciation lexicons for vowel-lengthening. We also compare GMM-HMM and DNN-HMM acoustic models. We obtained a remarkable improvement on lyric recognition in comparison with the baseline system for spontaneous speech recognition.
BibTeX
@inproceedings{icassp2016_speechanalysisof,
title = {Speech analysis of sung-speech and lyric recognition in monophonic singing},
author = {Dairoku Kawai and Kazumasa Yamamoto and Seiichi Nakagawa},
booktitle = {ICASSP 2016},
year = {2016}
}