← Search

Paola García

6 accepted papers

2023

Adapting Self-Supervised Models to Multi-Talker Speech Recognition Using Speaker Embeddings

ICASSP 2023accepted

Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-talker applications. However, these models often have degraded performance for multi-talker scenarios — possibly due to t…

Cited by 0SourceScholar
2023

Bridging Speech and Textual Pre-Trained Models With Unsupervised ASR

ICASSP 2023accepted

Spoken language understanding (SLU) is a task aiming to extract high-level semantics from spoken utterances. Previous works have investigated the use of speech self-supervised models and textual pre-trained models, which have shown reasonable improvements to various SLU tasks. However, because of th…

Cited by 0SourceScholar
2022

Investigating Self-Supervised Learning for Speech Enhancement and Separation

ICASSP 2022accepted

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target speech from interfering speakers. Despite a great number of supervised learning-based enhancement and separation methods ha…

Cited by 0SourceScholar
2022

Multi-Channel End-To-End Neural Diarization with Distributed Microphones

ICASSP 2022accepted

Recent progress on end-to-end neural diarization (EEND) has en-abled overlap-aware speaker diarization with a single neural net-work. This paper proposes to enhance EEND by using multi-channel signals from distributed microphones. We replace Transformer en-coders in EEND with two types of encoders t…

Cited by 0SourceScholar
2021

End-To-End Speaker Diarization as Post-Processing

ICASSP 2021accepted

This paper investigates the utilization of an end-to-end diarization model as post-processing of conventional clustering-based diarization. Clustering-based diarization methods partition frames into clusters of the number of speakers; thus, they typically cannot handle overlapping speech because eac…

Cited by 0SourceScholar
2020

Speaker Diarization with Region Proposal Network

ICASSP 2020accepted

Speaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in various scenarios, they are composed of several independently-optimized modules and c…

Cited by 0SourceScholar