ICASSP 2017accepted0 citations

Parallel phonetically aware DNNs and LSTM-RNNS for frame-by-frame discriminative modeling of spoken language identification

Ryo Masumura, Taichi Asami, Hirokazu Masataki, Yushi Aono

Abstract

Parallel phonetically aware deep neural networks (PPA-DNNs) and long short-term memory recurrent neural networks (PPA-LSTM-RNNs) to enhance frame-by-frame discriminative modeling of spoken language identification are proposed. This idea is inspired by traditional systems based on parallel phoneme recognition followed by language modeling (PPRLM). The proposed methods utilize multiple senone bottleneck features individually extracted from language-dependent senone-based DNNs in a frame-by-frame manner. The multiple senone bottleneck features can yield phonetic awareness to frame-by-frame DNNs and LSTM-RNNs without losing compatibility to real time applications. In experiments, three senone-based DNNs are introduced in order to extract senone bottleneck features, and both single use and parallel use of them are examined. Furthermore, we also examine a combination of PPA-DNNs and PPA-LSTM-RNNs. The proposed method's effectiveness is investigated by comparison with a simple speech aware modeling and traditional systems based on PPRLM.

BibTeX
@inproceedings{icassp2017_parallelphonetic,
  title = {Parallel phonetically aware DNNs and LSTM-RNNS for frame-by-frame discriminative modeling of spoken language identification},
  author = {Ryo Masumura and Taichi Asami and Hirokazu Masataki and Yushi Aono},
  booktitle = {ICASSP 2017},
  year = {2017}
}
Parallel phonetically aware DNNs and LSTM-RNNS for frame-by-frame discriminative modeling of spoken language identification · ICASSP 2017