ICASSP 2015accepted0 citations

Language-resource independent speech segmentation using cues from a spectrogram image

Su Jun Leow, Engsiong Chng, Chin-Hui Lee

Abstract

In this paper, we use image processing techniques on the speech spectrogram to perform speech phoneme segmentation. The proposed method relies solely on visual cues on the spectrogram, without the need for language-specific training data. The results are evaluated on the TIMIT corpus, and compared to other unsupervised speech segmentation techniques, with comparable results obtained. We also fuse the results with those obtained by hidden Markov models (HMM) and HMM-based forced alignment to investigate if image features can provide an additional feature representation for speech processing tasks. With the fusion, up to 10% absolute improvement in segmentation accuracy over the HMM baselines can be obtained. Results are promising and suggests a strong potential for image-based features applying to speech processing.

BibTeX
@inproceedings{icassp2015_languageresource,
  title = {Language-resource independent speech segmentation using cues from a spectrogram image},
  author = {Su Jun Leow and Engsiong Chng and Chin-Hui Lee},
  booktitle = {ICASSP 2015},
  year = {2015}
}