← Search

George Saon

26 accepted papers

2025

A Non-autoregressive Model for Joint STT and TTS

ICASSP 2025accepted

In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text modalities as input either individually or together. The proposed mo…

Cited by 0SourceScholar
2025

Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASR

ICASSP 2025accepted

There is an on-going body of research on training separate dedicated models for either short-form or long-form utterances. Multi-form acoustic models that are simply trained on combined data from long-form and short-form utterances often suffer from various negative impacts due to the diversity of a…

Cited by 0SourceScholar
2025

LLM based Text Generation for Improved Low-resource Speech Recognition Models

ICASSP 2025accepted

Limited transcribed spoken style data is a critical bottleneck in building automatic speech recognition (ASR) systems for low-resource languages. Prompting a large language model (LLM) to paraphrase input text can generate novel text data that is constrained to be semantically similar to the source…

Cited by 0SourceScholar
2024

Multiple Representation Transfer from Large Language Models to End-to-End ASR Systems

ICASSP 2024accepted

Transferring the knowledge of large language models (LLMs) is a promising technique to incorporate linguistic knowledge into end-to-end automatic speech recognition (ASR) systems. However, existing works only transfer a single representation of LLM (e.g. the last layer of pretrained BERT), while the…

Cited by 0SourceScholar
2023

Multi-Speaker Data Augmentation for Improved end-to-end Automatic Speech Recognition

ICASSP 2023accepted

Publicly available datasets traditionally used to train E2E ASR models for conversational telephone speech recognition are based on clean, short duration, single speaker utterances collected on separate channels. While E2E ASR models achieve state-of-the-art performance on recognition tasks that mat…

Cited by 0SourceScholar
2023

Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word Dictionaries

EMNLP 2023long main

Despite the impressive performance of ASR models on mainstream benchmarks, their performance on rare words is unsatisfactory. In enterprise settings, often a focused list of entities (such as locations, names, etc) are available which can be used to adapt the model to the terminology of specific dom…

Cited by 0SourceScholar
2022

Improving End-to-end Models for Set Prediction in Spoken Language Understanding

ICASSP 2022accepted

The goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts. Advances in end-to-end (E2E) speech modeling have made it possible to train solely on semantic entities, which are far…

Cited by 0SourceScholar
2022

Integrating Text Inputs for Training and Adapting RNN Transducer ASR Models

ICASSP 2022accepted

Compared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be in-dependently adapted to a new domain, recent end-to-end (E2E) ASR system are harder to customize due to their all-neural monolithic construction. In this paper, we propose a…

Cited by 0SourceScholar
2022

Speech Recognition Using Biologically-Inspired Neural Networks

ICASSP 2022accepted

Automatic speech recognition systems (ASR), such as the recurrent neural network transducer (RNN-T), have reached close to human-like performance and are deployed in commercial applications. However, their core operations depart from the powerful biological counterpart, the human brain. On the other…

Cited by 0SourceScholar
2022

Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding Systems

ICASSP 2022accepted

The lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly process speech inputs. In contrast, large amounts of text data with suitable labels are usually available. In this paper, we p…

Cited by 0SourceScholar
2021

Advancing RNN Transducer Technology for Speech Recognition

ICASSP 2021accepted

We investigate a set of techniques for RNN Transducers (RNN-Ts) that were instrumental in lowering the word error rate on three different tasks (Switchboard 300 hours, conversational Spanish 780 hours and conversational Italian 900 hours). The techniques pertain to architectural changes, speaker ada…

Cited by 0SourceScholar
2021

RNN Transducer Models for Spoken Language Understanding

ICASSP 2021accepted

We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available ann…

Cited by 0SourceScholar
2020

Improving Efficiency in Large-Scale Decentralized Distributed Training

ICASSP 2020accepted

Decentralized Parallel SGD (D-PSGD) and its asynchronous variant Asynchronous Parallel SGD (AD-PSGD) is a family of distributed learning algorithms that have been demonstrated to perform well for large-scale deep learning tasks. One drawback of (A)D-PSGD is that the spectral gap of the mixing matrix…

Cited by 0SourceScholar
2019

Distributed Deep Learning Strategies for Automatic Speech Recognition

ICASSP 2019accepted

In this paper, we propose and investigate a variety of distributed deep learning strategies for automatic speech recognition (ASR) and evaluate them with a state-of-the-art Long short-term memory (LSTM) acoustic model on the 2000-hour Switchboard (SWB2000), which is one of the most widely used datas…

Cited by 0SourceScholar
2019

English Broadcast News Speech Recognition by Humans and Machines

ICASSP 2019accepted

With recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broa…

Cited by 0SourceScholar
2019

Sequence Noise Injected Training for End-to-end Speech Recognition

ICASSP 2019accepted

We present a simple noise injection algorithm for training end-to-end ASR models which consists in adding to the spectra of training utterances the scaled spectra of random utterances of comparable length. We conjecture that the sequence information of the "noise" utterances is important and verify…

Cited by 0SourceScholar
2018

Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition

ICASSP 2018accepted

Direct acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional subword based automatic speech recognition models using phones, characters, or context-dependent hidden Markov model states. This is because A2W models recognize words from sp…

Cited by 0SourceScholar
2017

Knowledge distillation across ensembles of multilingual models for low-resource languages

ICASSP 2017accepted

This paper investigates the effectiveness of knowledge distillation in the context of multilingual models. We show that with knowledge distillation, Long Short-Term Memory(LSTM) models can be used to train standard feed-forward Deep Neural Network (DNN) models for a variety of low-resource languages…

Cited by 0SourceScholar
2017

Network architectures for multilingual speech representation learning

ICASSP 2017accepted

Multilingual (ML) representations play a key role in building speech recognition systems for low resource languages. The IARPA sponsored BABEL program focuses on building speech recognition (ASR) and keyword search (KWS) systems in over 24 languages with limited training data. The most common mechan…

Cited by 0SourceScholar
2016

On the importance of event detection for ASR

ICASSP 2016accepted

The performance of modern large vocabulary continuous speech recognition (LVCSR) systems is heavily affected by segment boundaries, proper speaker identification of the segments, as well as removal of spurious data. We propose to use Long Short Term Memory (LSTM) recurrent neural networks to partiti…

Cited by 0SourceScholar
2015

Improvements to the IBM speech activity detection system for the DARPA RATS program

ICASSP 2015accepted

In this paper we describe improvements to the IBM speech activity detection (SAD) system for the third phase of the DARPA RATS program. The progress during this final phase comes from jointly training convolutional and regular deep neural networks with rich time-frequency representations of speech.…

Cited by 93SourceScholar