← Search

Françoise Beaufays

21 accepted papers

2025

Speech Re-Painting for Robust ASR

ICASSP 2025accepted

Synthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synth…

Cited by 0SourceScholar
2024

Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data

ICASSP 2024accepted

Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-t…

Cited by 0SourceScholar
2024

Improving Speech Recognition for African American English with Audio Classification

ICASSP 2024accepted

Automatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to mitigate this is to train or fine-tune models with more representative datasets. But this approach can be hindered by lim…

Cited by 0SourceScholar
2023

Efficient Domain Adaptation for Speech Foundation Models

ICASSP 2023accepted

Foundation models (FMs), that are trained on broad data at scale and are adaptable to a wide range of downstream tasks, have brought large interest in the research community. Benefiting from the diverse data sources such as different modalities, languages and application domains, foundation models h…

Cited by 0SourceScholar
2023

Lego-Features: Exporting Modular Encoder Features for Streaming and Deliberation ASR

ICASSP 2023accepted

In end-to-end (E2E) speech recognition models, a representational tight-coupling inevitably emerges between the encoder and the decoder. We build upon recent work that has begun to explore building encoders with modular encoded representations, such that encoders and decoders from different models c…

Cited by 3SourceScholar
2023

Online Model Compression for Federated Learning with Large Models

ICASSP 2023accepted

This paper addresses the challenges of training large neural networks under federated learning settings: high on-device memory usage and communication cost. The proposed Online Model Compression (OMC) provides a framework that stores model parameters in a compressed format and decompresses them only…

Cited by 0SourceScholar
2022

A Method to Reveal Speaker Identity in Distributed ASR Training, and How to Counter IT

ICASSP 2022accepted

End-to-end Automatic Speech Recognition (ASR) models are commonly trained over spoken utterances using optimization methods like Stochastic Gradient Descent (SGD). In distributed settings like Federated Learning, model training requires transmission of gradients over a network. In this work, we desi…

Cited by 0SourceScholar
2022

Enabling On-Device Training of Speech Recognition Models With Federated Dropout

ICASSP 2022accepted

Federated learning can be used to train machine learning models on the edge on local data that never leave devices, providing privacy by default. This presents a challenge pertaining to the communication and computation costs associated with clients’ devices. These costs are strongly correlated with…

Cited by 0SourceScholar
2022

Exploring Heterogeneous Characteristics of Layers in ASR Models for More Efficient Training

ICASSP 2022accepted

Transformer-based architectures have been the subject of research aimed at understanding their overparameterization and the non-uniform importance of their layers. Applying these approaches to Automatic Speech Recognition, we demonstrate that the state-of-the-art Conformer models generally have mult…

Cited by 0SourceScholar
2022

Fast Contextual Adaptation with Neural Associative Memory for On-Device Personalized Speech Recognition

ICASSP 2022accepted

Fast contextual adaptation has shown to be effective in improving Automatic Speech Recognition (ASR) of rare words and when combined with an on-device personalized training, it can yield an even better recognition result. However, the traditional re-scoring approaches based on an external language m…

Cited by 0SourceScholar
2022

Large-Scale ASR Domain Adaptation Using Self- and Semi-Supervised Learning

ICASSP 2022accepted

Self- and semi-supervised learning methods have been actively investigated to reduce labeled training data or enhance model performance. However, these approaches mostly focus on in-domain performance for public datasets. In this study, we utilize the combination of self- and semi-supervised learnin…

Cited by 0SourceScholar
2022

Partial Variable Training for Efficient on-Device Federated Learning

ICASSP 2022accepted

This paper aims to address the major challenges of Federated Learning (FL) on edge devices: limited memory and expensive communication. We propose a novel method, called Partial Variable Training (PVT), that only trains a small subset of variables on edge devices to reduce memory usage and communica…

Cited by 0SourceScholar
2021

Revealing and Protecting Labels in Distributed Training

NeurIPS 2021poster

Distributed learning paradigms such as federated learning often involve transmission of model updates, or gradients, over a network, thereby avoiding transmission of private data. However, it is possible for sensitive information about the training data to be revealed from such gradients. Prior work…

2021

Training Speech Recognition Models with Federated Learning: A Quality/Cost Framework

ICASSP 2021accepted

We propose using federated learning, a decentralized on-device learning paradigm, to train speech recognition models. By performing epochs of training on a per-user basis, federated learning must incur the cost of dealing with non-IID data distributions, which are expected to negatively affect the q…

Cited by 0SourceScholar
2020

Low-Rank Gradient Approximation for Memory-Efficient on-Device Training of Deep Neural Network

ICASSP 2020accepted

Training machine learning models on mobile devices has the potential of improving both privacy and accuracy of the models. However, one of the major obstacles to achieving this goal is the memory limitation of mobile devices. Reducing training memory enables models with high-dimensional weight matri…

Cited by 0SourceScholar
2016

Personalized speech recognition on mobile devices

ICASSP 2016accepted

We describe a large vocabulary speech recognition system that is accurate, has low latency, and yet has a small enough memory and computational footprint to run faster than real-time on a Nexus 5 Android smartphone. We employ a quantized Long Short-Term Memory (LSTM) acoustic model trained with conn…

Cited by 0SourceScholar
2015

Fix it where it fails: Pronunciation learning by mining error corrections from speech logs

ICASSP 2015accepted

The pronunciation dictionary, or lexicon, is an essential component in an automatic speech recognition (ASR) system in that incorrect pronunciations cause systematic misrecognitions. It typically consists of a list of word-pronunciation pairs written by linguists, and a grapheme-to-phoneme (G2P) eng…

Cited by 0SourceScholar
2015

Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networks

ICASSP 2015accepted

Grapheme-to-phoneme (G2P) models are key components in speech recognition and text-to-speech systems as they describe how words are pronounced. We propose a G2P model based on a Long Short-Term Memory (LSTM) recurrent neural network (RNN). In contrast to traditional joint-sequence based G2P approach…

Cited by 0SourceScholar
2015

Learning acoustic frame labeling for speech recognition with recurrent neural networks

ICASSP 2015accepted

We explore alternative acoustic modeling techniques for large vocabulary speech recognition using Long Short-Term Memory recurrent neural networks. For an acoustic frame labeling task, we compare the conventional approach of cross-entropy (CE) training using fixed forced-alignments of frames and lab…

Cited by 0SourceScholar
2015

Long short term memory neural network for keyboard gesture decoding

ICASSP 2015accepted

Gesture typing is an efficient input method for phones and tablets using continuous traces created by a pointed object (e.g., finger or stylus). Translating such continuous gestures into textual input is a challenging task as gesture inputs exhibit many features found in speech and handwriting such…

Cited by 0SourceScholar