ICASSP 2018accepted0 citations

End-To-End Low-Resource Lip-Reading with Maxout Cnn and Lstm

Ivan Fung, Brian Mak

Abstract

Lip-reading is the task of recognizing speech solely from the visual movement of the mouth. Although recent works have demonstrated the effectiveness of convolutional neural network (CNN) and long short-term memory (LSTM) recurrent neural network in lip-reading, similar architectures under low-resource scenario have not yet been explored. Our proposed end - to-end deep learning model fuses conventional CNN and bidirectional LSTM (BLSTM) together with max-out activation units (maxout-CNN-BLSTM), and is capable of attaining a word accuracy of 87.6% on the Ouluvs2 corpus, offering an absolute improvement of 3.1 % to the previous state-of-the-art auto-encoder-BLSTM model. To the best of our knowledge, this is the first end - to-end low -resource lip-reading system that does not require any separate feature extraction stage nor pre-training phase with external data resources. This is also the first work that utilizes maxout units in both CNN and LSTM in one single deep neural network.

BibTeX
@inproceedings{icassp2018_endtoendlowresou,
  title = {End-To-End Low-Resource Lip-Reading with Maxout Cnn and Lstm},
  author = {Ivan Fung and Brian Mak},
  booktitle = {ICASSP 2018},
  year = {2018}
}
End-To-End Low-Resource Lip-Reading with Maxout Cnn and Lstm · ICASSP 2018