← Search

Leda Sari

11 accepted papers

2025

CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR

ICASSP 2025accepted

CTC compressor can be an effective approach to integrate audio encoders to decoder-only models, which has gained growing interest for different speech applications. In this work, we propose a novel CTC compressor based joint speech and text training (CJST) framework for decoder-only ASR. CJST matche…

Cited by 0SourceScholar
2024

Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of a Multilingual ASR Model

ICASSP 2024accepted

Neural network pruning offers an effective method for compressing a multilingual automatic speech recognition (ASR) model with minimal performance loss. However, it entails several rounds of pruning and re-training needed to be run for each language. In this work, we propose the use of an adaptive m…

Cited by 0SourceScholar
2023

Self-Supervised Representations for Singing Voice Conversion

ICASSP 2023accepted

A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Vec 2.0 have helped further the state-of-the-art. Though these methods produce mor…

Cited by 25SourceScholar
2023

Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

NeurIPS 2023poster

Large-scale generative models such as GPT and DALL-E have revolutionized the research community. These models not only generate high fidelity outputs, but are also generalists which can solve tasks not explicitly taught. In contrast, speech generative models are still primitive in terms of scale and…

Cited by 299SourcePDFScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions

ICASSP 2022accepted

The problem of machine learning systems demonstrating bias towards specific groups of individuals has been studied extensively, particularly in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR). This paper presents initial Speech Recognition results on “Casual Conve…

Cited by 53SourceScholar
2021

A Multi-View Approach to Audio-Visual Speaker Verification

ICASSP 2021accepted

Although speaker verification has conventionally been an audio-only task, some practical applications provide both audio and visual streams of input. In these cases, the visual stream provides complementary information and can often be leveraged in conjunction with the acoustics of speech to improve…

Cited by 0SourceScholar
2021

Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question Answering

NAACL 2021long

Although Question-Answering has long been of research interest, its accessibility to users through a speech interface and its support to multiple languages have not been addressed in prior studies. Towards these ends, we present a new task and a synthetically-generated dataset to do Fact-based Visua…

2020

Training Spoken Language Understanding Systems with Non-Parallel Speech and Text

ICASSP 2020accepted

End-to-end spoken language understanding (SLU) systems are typically trained on large amounts of data. In many practical scenarios, the amount of labeled speech is often limited as opposed to text. In this study, we investigate the use of non-parallel speech and text to improve the performance of di…

Cited by 0SourceScholar
2020

Unsupervised Speaker Adaptation Using Attention-Based Speaker Memory for End-to-End ASR

ICASSP 2020accepted

We propose an unsupervised speaker adaptation method inspired by the neural Turing machine for end-to-end (E2E) automatic speech recognition (ASR). The proposed model contains a memory block that holds speaker i-vectors extracted from the training data and reads relevant i-vectors from the memory th…

Cited by 0SourceScholar
2019

Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast News

ICASSP 2019accepted

In this work, we investigate pre-training of neural network based speaker embeddings for low-latency speaker change detection. Our proposed system takes two speech segments, generates embeddings using shared Siamese layers and then classifies the concatenated embeddings depending on whether they are…

Cited by 0SourceScholar