← Search

William Hartmann

9 accepted papers

2022

Combining Unsupervised and Text Augmented Semi-Supervised Learning For Low Resourced Autoregressive Speech Recognition

ICASSP 2022accepted

Recent advances in unsupervised representation learning have demonstrated the impact of pretraining on large amounts of read speech. We adapt these techniques for domain adaptation in low-resource—both in terms of data and compute—conversational and broadcast domains. Moving beyond CTC, we pretrain…

Cited by 0SourceScholar
2021

Improved Data Selection for Domain Adaptation in ASR

ICASSP 2021accepted

Automatic speech recognition (ASR) systems are highly sensitive to train-test domain mismatch. However, because transcription is often prohibitively expensive, it is important to be able to make use of available transcribed out-of-domain data. We address the problem of domain adaptation with semi-su…

Cited by 0SourceScholar
2020

Towards a New Understanding of the Training of Neural Networks with Mislabeled Training Data

ICASSP 2020accepted

We investigate the problem of machine learning with mislabeled training data. We try to make the effects of mislabeled training better understood through analysis of the basic model and equations that characterize the problem. This includes results about the ability of the noisy model to make the sa…

Cited by 0SourceScholar
2019

Learning from the Best: A Teacher-student Multilingual Framework for Low-resource Languages

ICASSP 2019accepted

The traditional method of pretraining neural acoustic models in low-resource languages consists of initializing the acoustic model parameters with a large, annotated multilingual corpus and can be a drain on time and resources. In an attempt to reuse TDNN-LSTMs already pre-trained using multilingual…

Cited by 0SourceScholar
2018

Individual Ship Detection Using Underwater Acoustics

ICASSP 2018accepted

Individual ship detection from underwater audio is the task of deciding whether a specific ship is present, using sound captured by an underwater hydrophone. It is a task analogous to speaker identification (SID), in the sense that it is an open-class detection task; the ships present could be other…

Cited by 0SourceScholar
2018

Optimizing Multilingual Knowledge Transfer for Time-Delay Neural Networks with Low-Rank Factorization

ICASSP 2018accepted

When producing speech-to-text (STT) systems on a lower resource language, it is often beneficial to use knowledge obtained from a significantly larger multilingual dataset. We have seen benefits from using a multilingual TDNN as initialization for training an acoustic model on a target low resource…

Cited by 0SourceScholar
2017

Analysis of keyword spotting performance across IARPA babel languages

ICASSP 2017accepted

With the completion of the IARPA Babel program, it is possible to systematically analyze the performance of speech recognition systems across a wide variety of languages. We select 16 languages from the dataset and compare performance using a deep neural network-based acoustic model. The focus is on…

Cited by 0SourceScholar
2017

The 2016 BBN Georgian telephone speech keyword spotting system

ICASSP 2017accepted

In this paper we describe the 2016 BBN conversational telephone speech keyword spotting system; the culmination of four years of research and development under the IARPA Babel program. The system was constructed in response to the NIST Open Keyword Search (OpenKWS) evaluation of 2016. We present our…

Cited by 0SourceScholar