← Search

Jasha Droppo

18 accepted papers

2023

Federated Self-Learning with Weak Supervision for Speech Recognition

ICASSP 2023accepted

Automatic speech recognition (ASR) models with low-footprint are increasingly being deployed on edge devices for conversational agents, which enhances privacy. We study the problem of federated continual incremental learning for recurrent neural network-transducer (RNN-T) ASR models in the privacy-e…

Cited by 0SourceScholar
2022

Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology

ICASSP 2022accepted

Acoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is…

Cited by 0SourceScholar
2022

Improving Fairness in Speaker Verification via Group-Adapted Fusion Network

ICASSP 2022accepted

Modern speaker verification models use deep neural networks to encode utterance audio into discriminative embedding vectors. During the training process, these networks are typically optimized to differentiate arbitrary speakers. This learning process biases the learning of fine voice characteristic…

Cited by 0SourceScholar
2021

DO as I Mean, Not as I Say: Sequence Loss Training for Spoken Language Understanding

ICASSP 2021accepted

Spoken language understanding (SLU) systems extract transcriptions, as well as semantics of intent or named entities from speech, and are essential components of voice activated systems. SLU models, which either directly extract semantics from audio or are composed of pipelined automatic speech reco…

Cited by 0SourceScholar
2021

Exploring the application of synthetic audio in training keyword spotters

ICASSP 2021accepted

The study of keyword spotting, a subfield within the broader field of speech recognition that centers around identifying individual keywords in speech audio, has gained particular importance in recent years with the rise of personal voice assistants such as Alexa. As voice assistants aim to rapidly…

Cited by 0SourceScholar
2021

Joint ASR and Language Identification Using RNN-T: An Efficient Approach to Dynamic Language Switching

ICASSP 2021accepted

Conventional dynamic language switching enables seamless multilingual interactions by running several monolingual ASR systems in parallel and triggering the appropriate downstream components using a standalone language identification (LID) service. Since this solution is neither scalable nor cost- a…

Cited by 0SourceScholar
2021

Top-Down Attention in End-to-End Spoken Language Understanding

ICASSP 2021accepted

Spoken language understanding (SLU) is the task of inferring the semantics of spoken utterances. Traditionally, this has been achieved with a cascading combination of Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU) modules that are optimized separately, which can lead to…

Cited by 0SourceScholar
2019

Single-channel Speech Extraction Using Speaker Inventory and Attention Network

ICASSP 2019accepted

Neural network-based speech separation has received a surge of interest in recent years. Previously proposed methods either are speaker independent or extract a target speaker's voice by using his or her voice snippet. In applications such as home devices or office meeting transcriptions, a possible…

Cited by 0SourceScholar
2018

The Microsoft 2017 Conversational Speech Recognition System

ICASSP 2018accepted

We describe the latest version of Microsoft's conversational speech recognition system for the Switchboard and CallHome domains. The system adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language mo…

Cited by 0SourceScholar
2017

The microsoft 2016 conversational speech recognition system

ICASSP 2017accepted

We describe Microsoft's conversational speech recognition system, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard recognition task. Inspired by machine learning ensemble techniques, the system uses a ra…

Cited by 0SourceScholar
2016

Parallelizing WFST speech decoders

ICASSP 2016accepted

The performance-intensive part of a large-vocabulary continuous speech-recognition system is the Viterbi computation that determines the sequence of words that are most likely to generate the acoustic-state scores extracted from an input utterance. This paper presents an efficient parallel algorithm…

Cited by 0SourceScholar
2015

Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learning

ICASSP 2015accepted

In this paper, we propose two approaches to improve deep neural network (DNN) acoustic models for speech recognition in reverberant environments. Both methods utilize auxiliary information in training the DNN but differ in the type of information and the manner in which it is used. The first method…

Cited by 0SourceScholar
2015

Speech recognition with prediction-adaptation-correction recurrent neural networks

ICASSP 2015accepted

We propose the prediction-adaptation-correction RNN (PAC-RNN), in which a correction DNN estimates the state posterior probability based on both the current frame and the prediction made on the past frames by a prediction DNN. The result from the main DNN is fed back to the prediction DNN to make be…

Cited by 0SourceScholar