← Search

Hank Liao

5 accepted papers

2024

Conformer is All You Need for Visual Speech Recognition

ICASSP 2024accepted

Visual speech recognition models extract visual features in a hierarchical manner. At the lower level, there is a visual front-end with a limited temporal receptive field that processes the raw pixels depicting the lips or faces. At the higher level, there is an encoder that attends to the embedding…

Cited by 0SourceScholar
2024

USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models

ICASSP 2024accepted

We introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a speech foundation model trained on a large quantity of supervised and unsupervised data, demonstrating the utility of fine-…

Cited by 0SourceScholar
2020

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition

ICASSP 2020accepted

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are potentially on screen one needs to decide which face to feed to the…

Cited by 0SourceScholar
2018

RADMM: Recurrent Adaptive Mixture Model with Applications to Domain Robust Language Modeling

ICASSP 2018accepted

We present a new architecture and a training strategy for an adaptive mixture of experts with applications to domain robust language modeling. The proposed model is designed to benefit from the scenario where the training data are available in diverse domains as is the case for YouTube speech recogn…

Cited by 0SourceScholar
2015

Exemplar-based large vocabulary speech recognition using k-nearest neighbors

ICASSP 2015accepted

This paper describes a large scale exemplar-based acoustic modeling approach for large vocabulary continuous speech recognition. We construct an index of labeled training frames using high-level features extracted from the bottleneck layer of a deep neural network as indexing features. At recognitio…

Cited by 0SourceScholar