← Search

Pedro J. Moreno

14 accepted papers

2024

Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models

ICASSP 2024accepted

The accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. Widespread deployment and adoption of these models, however, requires computationally efficient strategies for decoding.…

Cited by 11SourceScholar
2023

Large-Scale Language Model Rescoring on Long-Form Data

ICASSP 2023accepted

In this work, we study the impact of Large-scale Language Models (LLM) on Automated Speech Recognition (ASR) of YouTube videos, which we use as a source for long-form ASR. We demonstrate up to 8% relative reduction in Word Error Eate (WER) on US English (en-us) and code-switched Indian English (en-i…

Cited by 27SourceScholar
2023

Modular Conformer Training for Flexible End-to-End ASR

ICASSP 2023accepted

The state-of-the-art conformer used in automatic speech recognition combines feed-forward, convolution and multi-headed self-attention layers in a single model that is trained end-to-end with a decoder network. While this end-to-end training is simple and beneficial for word error rate, it restricts…

Cited by 0SourceScholar
2022

Multilingual Second-Pass Rescoring for Automatic Speech Recognition Systems

ICASSP 2022accepted

Second-pass rescoring is a well known technique to improve the performance of Automatic Speech Recognition (ASR) systems. Neural Oracle Search (NOS), which selects the most likely hypothesis from an N-best hypothesis list by integrating information from multiple sources, such as the input acoustic r…

Cited by 0SourceScholar
2022

Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive Losses

ICASSP 2022accepted

An effective way to learn representations from untranscribed speech and unspoken text with linguistic/lexical representations derived from synthesized speech was introduced in tts4pretrain [1]. However, the representations learned from synthesized and real speech are likely to be different, potentia…

Cited by 0SourceScholar
2021

Extending Parrotron: An End-to-End, Speech Conversion and Speech Recognition Model for Atypical Speech

ICASSP 2021accepted

We present an extended Parrotron model: a single, end-to-end network that enables voice conversion and recognition simultaneously. Input spectrograms are transformed to output spectrograms in the voice of a predetermined target speaker while also generating hypotheses in a target vocabulary. We stud…

Cited by 0SourceScholar
2021

Mixture of Informed Experts for Multilingual Speech Recognition

ICASSP 2021accepted

When trained on related or low-resource languages, multilingual speech recognition models often outperform their monolingual counterparts. However, these models can suffer from loss in performance for high resource or unrelated languages. We investigate the use of a mixture-of-experts approach to as…

Cited by 48SourceScholar
2020

Improving Speech Recognition Using Consistent Predictions on Synthesized Speech

ICASSP 2020accepted

Speech synthesis has advanced to the point of being close to indistinguishable from human speech. However, efforts to train speech recognition systems on synthesized utterances have not been able to show that synthesized data can be effectively used to augment or replace human speech. In this work,…

Cited by 0SourceScholar
2020

Neural Oracle Search on N-BEST Hypotheses

ICASSP 2020accepted

In this paper, we propose a neural search algorithm to select the most likely hypothesis using a sequence of acoustic representations and multiple hypotheses as input. The algorithm provides a sequence level score for each audio-hypothesis pair that is obtained by integrating information from multip…

Cited by 0SourceScholar
2018

Modeling Non-Linguistic Contextual Signals in LSTM Language Models Via Domain Adaptation

ICASSP 2018accepted

Language Models (LMs) for Automatic Speech Recognition (ASR) can benefit from utilizing non-linguistic contextual signals in modeling. Examples of these signals include the geographical location of the user speaking to the system and/or the identity of the application (app) being spoken to. In pract…

Cited by 0SourceScholar
2018

Multilingual Speech Recognition with a Single End-to-End Model

ICASSP 2018accepted

Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual ASR because they encapsula…

Cited by 292SourceScholar
2016

Selection and combination of hypotheses for dialectal speech recognition

ICASSP 2016accepted

While research has often shown that building dialect-specific Automatic Speech Recognizers is the optimal approach to dealing with dialectal variations of the same language, we have observed that dialect-specific recognizers do not always output the best recognitions. Often enough, another dialectal…

Cited by 15SourceScholar
2015

Improved recognition of contact names in voice commands

ICASSP 2015accepted

The recognition of contact names in mobile-device voice commands is a challenging problem. Some of the difficulties include potentially infinite vocabularies, low probability of contact tokens in the language model (LM), increased false triggering of contact voice commands when none are spoken, and…

Cited by 0SourceScholar