← Search

Catalin Zorila

9 accepted papers

2024

Geodesic Interpolation of Frame-Wise Speaker Embeddings for the Diarization of Meeting Scenarios

ICASSP 2024accepted

We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geodesic distance loss is used that enforces the embeddings computed from regions w…

Cited by 0SourceScholar
2023

Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting Diarization

ICASSP 2023accepted

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings…

Cited by 0SourceScholar
2023

On the Effectiveness of Monoaural Target Source Extraction for Distant end-to-end Automatic Speech Recognition

ICASSP 2023accepted

Recent work on enhancement has shown that frequency domain methods may outperform the time domain approaches, while most of the prior art is focused on reporting objective enhancement metrics on simulated noisy data or use less modern hybrid acoustic models for evaluation. In this paper we investiga…

Cited by 0SourceScholar
2022

Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech Recognition

ICASSP 2022accepted

Improving the accuracy of single-channel automatic speech recognition (ASR) in noisy conditions is challenging. Strong speech enhancement front-ends are available, however, they typically require that the ASR model is retrained to cope with the processing artifacts. In this paper we explore a speake…

Cited by 0SourceScholar
2021

Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism

ICASSP 2021accepted

In this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on an improved multi-channel time-domain speech separation network which employs spe…

Cited by 0SourceScholar
2020

On End-to-end Multi-channel Time Domain Speech Separation in Reverberant Environments

ICASSP 2020accepted

This paper introduces a new method for multi-channel time domain speech separation in reverberant environments. A fully-convolutional neural network structure has been used to directly separate speech from multiple microphone recordings, with no need of conventional spatial feature extraction. To re…

Cited by 0SourceScholar