← Search

Ladislav Mosner

8 accepted papers

2025

CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification

ICASSP 2025accepted

Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-hea…

Cited by 0SourceScholar
2023

Parameter-Efficient Transfer Learning of Pre-Trained Transformer Models for Speaker Verification Using Adapters

ICASSP 2023accepted

Recently, the pre-trained Transformer models have received a rising interest in the field of speech processing thanks to their great success in various downstream tasks. However, most fine-tuning approaches update all the parameters of the pre-trained model, which becomes prohibitive as the model si…

Cited by 0SourceScholar
2023

Speech-Based Emotion Recognition with Self-Supervised Models Using Attentive Channel-Wise Correlations and Label Smoothing

ICASSP 2023accepted

When recognizing emotions from speech, we encounter two common problems: how to optimally capture emotion-relevant information from the speech signal and how to best quantify or categorize the noisy subjective emotion labels. Self-supervised pre-trained representations can robustly capture informati…

Cited by 0SourceScholar
2022

Multi-Channel Speaker Verification with Conv-Tasnet Based Beamformer

ICASSP 2022accepted

We focus on the problem of speaker recognition in far-field multichannel data. The main contribution is introducing an alternative way of predicting spatial covariance matrices (SCMs) for a beamformer from the time domain signal. We propose to use ConvTasNet, a well-known source separation model, an…

Cited by 0SourceScholar
2022

Multisv: Dataset for Far-Field Multi-Channel Speaker Verification

ICASSP 2022accepted

Motivated by unconsolidated data situation and the lack of a standard benchmark in the field, we complement our previous efforts and present a comprehensive corpus designed for training and evaluating text-independent multi-channel speaker verification systems. It can be readily used also for experi…

Cited by 0SourceScholar
2020

But System for the Second Dihard Speech Diarization Challenge

ICASSP 2020accepted

This paper describes the winning systems developed by the BUT team for the four tracks of the Second DIHARD Speech Diarization Challenge. For tracks 1 and 2 the systems were mainly based on performing agglomerative hierarchical clustering (AHC) of x-vectors, followed by another x-vector clustering b…

Cited by 0SourceScholar
2019

Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning

ICASSP 2019accepted

For real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we a…

Cited by 0SourceScholar
2018

Dereverberation and Beamforming in Far-Field Speaker Recognition

ICASSP 2018accepted

This paper deals with far-field speaker recognition. On a corpus of NIST SRE 2010 data retransmitted in a real room with multiple microphones, we first demonstrate how room acoustics cause significant degradation of state-of-the-art i-vector based speaker recognition system. We then investigate seve…

Cited by 0SourceScholar