← Search

Michael Picheny

16 accepted papers

2023

A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at Scale

ICASSP 2023accepted

Unpaired text and audio injection have emerged as dominant methods for improving ASR performance in the absence of a large labeled corpus. However, little guidance exists on deploying these methods to improve production ASR systems that are trained on very large supervised corpora and with realistic…

Cited by 0SourceScholar
2022

Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions

ICASSP 2022accepted

The problem of machine learning systems demonstrating bias towards specific groups of individuals has been studied extensively, particularly in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR). This paper presents initial Speech Recognition results on “Casual Conve…

Cited by 53SourceScholar
2021

Multimodal Clustering Networks for Self-Supervised Learning From Unlabeled Videos

ICCV 2021poster

Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a framework that, starting from a pre-trained backbone,…

Cited by 110PDFcodeScholar
2020

Improving Efficiency in Large-Scale Decentralized Distributed Training

ICASSP 2020accepted

Decentralized Parallel SGD (D-PSGD) and its asynchronous variant Asynchronous Parallel SGD (AD-PSGD) is a family of distributed learning algorithms that have been demonstrated to perform well for large-scale deep learning tasks. One drawback of (A)D-PSGD is that the spectral gap of the mixing matrix…

Cited by 0SourceScholar
2020

Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent Systems

ICASSP 2020accepted

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can…

Cited by 0SourceScholar
2019

Acoustically Grounded Word Embeddings for Improved Acoustics-to-word Speech Recognition

ICASSP 2019accepted

Direct acoustics-to-word (A2W) systems for end-to-end automatic speech recognition are simpler to train, and more efficient to decode with, than sub-word systems. However, A2W systems can have difficulties at training time when data is limited, and at decoding time when recognizing words outside the…

Cited by 0SourceScholar
2019

Distributed Deep Learning Strategies for Automatic Speech Recognition

ICASSP 2019accepted

In this paper, we propose and investigate a variety of distributed deep learning strategies for automatic speech recognition (ASR) and evaluate them with a state-of-the-art Long short-term memory (LSTM) acoustic model on the 2000-hour Switchboard (SWB2000), which is one of the most widely used datas…

Cited by 0SourceScholar
2019

English Broadcast News Speech Recognition by Humans and Machines

ICASSP 2019accepted

With recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broa…

Cited by 0SourceScholar
2019

Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast News

ICASSP 2019accepted

In this work, we investigate pre-training of neural network based speaker embeddings for low-latency speaker change detection. Our proposed system takes two speech segments, generates embeddings using shared Siamese layers and then classifies the concatenated embeddings depending on whether they are…

Cited by 0SourceScholar
2018

Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition

ICASSP 2018accepted

Direct acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional subword based automatic speech recognition models using phones, characters, or context-dependent hidden Markov model states. This is because A2W models recognize words from sp…

Cited by 0SourceScholar
2018

Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural Networks

NeurIPS 2018poster

We propose a population-based Evolutionary Stochastic Gradient Descent (ESGD) framework for optimizing deep neural networks. ESGD combines SGD and gradient-free evolutionary algorithms as complementary algorithms in one framework in which the optimization alternates between the SGD step and evolutio…

2017

End-to-end speech recognition and keyword search on low-resource languages

ICASSP 2017accepted

In recent years, so-called, “end-to-end” speech recognition systems have emerged as viable alternatives to traditional ASR frameworks. Keyword search, localizing an orthographic query in a speech corpus, is typically performed by using automatic speech recognition (ASR) to generate an index. Previou…

Cited by 58SourceScholar
2017

Training variance and performance evaluation of neural networks in speech

ICASSP 2017accepted

In this work we study variance in the results of neural network training on a wide variety of configurations in automatic speech recognition. Although this variance itself is well known, this is, to the best of our knowledge, the first paper that performs an extensive empirical study on its effects…

Cited by 0SourceScholar
2016

A comparison between deep neural nets and kernel acoustic models for speech recognition

ICASSP 2016accepted

We study large-scale kernel methods for acoustic modeling and compare to DNNs on performance metrics related to both acoustic modeling and recognition. Measuring perplexity and frame-level classification accuracy, kernel-based acoustic models are as effective as their DNN counterparts. However, on t…

Cited by 0SourceScholar
2016

On the importance of event detection for ASR

ICASSP 2016accepted

The performance of modern large vocabulary continuous speech recognition (LVCSR) systems is heavily affected by segment boundaries, proper speaker identification of the segments, as well as removal of spurious data. We propose to use Long Short Term Memory (LSTM) recurrent neural networks to partiti…

Cited by 0SourceScholar