← Search

Man-Hung Siu

6 accepted papers

2024

Personalization of CTC-Based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization

ICASSP 2024accepted

Recent advances in deep learning and automatic speech recognition have improved the accuracy of end-to-end speech recognition systems, but recognition of personal content such as contact names remains a challenge. In this work, we describe our personalization solution for an end-to-end speech recogn…

Cited by 0SourceScholar
2020

Towards a New Understanding of the Training of Neural Networks with Mislabeled Training Data

ICASSP 2020accepted

We investigate the problem of machine learning with mislabeled training data. We try to make the effects of mislabeled training better understood through analysis of the basic model and equations that characterize the problem. This includes results about the ability of the noisy model to make the sa…

Cited by 0SourceScholar
2018

Optimizing Multilingual Knowledge Transfer for Time-Delay Neural Networks with Low-Rank Factorization

ICASSP 2018accepted

When producing speech-to-text (STT) systems on a lower resource language, it is often beneficial to use knowledge obtained from a significantly larger multilingual dataset. We have seen benefits from using a multilingual TDNN as initialization for training an acoustic model on a target low resource…

Cited by 0SourceScholar
2017

Unsupervised adaptation for deep neural networks using Alternating Direction Method of Multipliers

ICASSP 2017accepted

In this paper, we continue our work on linear least squares based adaptation (LLS) for deep neural networks. We show that our previously proposed algorithm is a special case of an optimization algorithm called Alternating Direction Method of Multipliers (ADMM). We demonstrate that the adaptation alg…

Cited by 0SourceScholar
2016

Importance sampling of delta-AUC: A basis for active learning for improved keyword search

ICASSP 2016accepted

We present an importance sampling based approach to the active learning problem of selecting additional training data to supplement a seed model. Our proposed Δ-AUC selection optimizes AUC improvement in keyword search and is evaluated on the Spanish Fisher corpus. We show that over different traini…

Cited by 0SourceScholar
2015

Large-scale speaker search using PLDA on mismatched conditions

ICASSP 2015accepted

Recent work reported on fast speaker search over large speech data corpora has focused on using locality sensitive hashing (LSH) search with hashing functions approximating i-vector based cosine distances (CosDist) for model comparisons. Because of the superior performance of probabilistic linear di…

Cited by 0SourceScholar