ICASSP 2016accepted0 citations

Improved DNN-based segmentation for multi-genre broadcast audio

Linlin Wang, Chao Zhang, Philip C. Woodland, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin, Xunying Liu, Yanmin Qian

Abstract

Automatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for multi-genre broadcast audio with deep neural network (DNN) based speech/non-speech detection. A further stage of change-point detection and clustering is used to obtain homogeneous segments. Suitable DNN inputs, context window sizes and architectures are studied with a series of experiments using a large corpus of MGB television audio. For MGB transcription, the improved segmenter yields roughly half the increase in word error rate, over manual segmentation, compared to the baseline DNN segmenter supplied for the 2015 ASRU MGB challenge.

BibTeX
@inproceedings{icassp2016_improveddnnbased,
  title = {Improved DNN-based segmentation for multi-genre broadcast audio},
  author = {Linlin Wang and Chao Zhang and Philip C. Woodland and Mark J. F. Gales and Panagiota Karanasou and Pierre Lanchantin and Xunying Liu and Yanmin Qian},
  booktitle = {ICASSP 2016},
  year = {2016}
}
Improved DNN-based segmentation for multi-genre broadcast audio · ICASSP 2016