← Search

Jan Cernocký

16 accepted papers

2025

CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification

ICASSP 2025accepted

Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-hea…

Cited by 0SourceScholar
2025

TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models

ICASSP 2025accepted

Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in nois…

Cited by 0SourceScholar
2025

Target Speaker ASR with Whisper

ICASSP 2025accepted

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs than to learn th…

Cited by 0SourceScholar
2024

Diacorrect: Error Correction Back-End for Speaker Diarization

ICASSP 2024accepted

In this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way. This method is inspired by error correction techniques in automatic speech recognition. Our model consists of two parallel convolutional encoders and…

Cited by 0SourceScholar
2024

Target Speech Extraction with Pre-Trained Self-Supervised Learning Models

ICASSP 2024accepted

Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been fully exploited. TSE aims to extract the speech of a target speaker in a mixture guided by enrollment utterances. We exp…

Cited by 0SourceScholar
2023

Parameter-Efficient Transfer Learning of Pre-Trained Transformer Models for Speaker Verification Using Adapters

ICASSP 2023accepted

Recently, the pre-trained Transformer models have received a rising interest in the field of speech processing thanks to their great success in various downstream tasks. However, most fine-tuning approaches update all the parameters of the pre-trained model, which becomes prohibitive as the model si…

Cited by 0SourceScholar
2022

DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation and Extraction

ICASSP 2022accepted

In recent years, a number of time-domain speech separation methods have been proposed. However, most of them are very sensitive to the environments and wide domain coverage tasks. In this paper, from the time-frequency domain perspective, we propose a densely-connected pyramid complex convolutional…

Cited by 0SourceScholar
2021

A Hierarchical Subspace Model for Language-Attuned Acoustic Unit Discovery

ICASSP 2021accepted

In this work, we propose a hierarchical subspace model for acoustic unit discovery. In this approach, we frame the task as one of learning embeddings on a low-dimensional phonetic subspace, and simultaneously specify the subspace itself as an embedding on a hyper-subspace. We train the hyper-subspac…

Cited by 12SourceScholar
2020

Investigation of Specaugment for Deep Speaker Embedding Learning

ICASSP 2020accepted

SpecAugment is a newly proposed data augmentation method for speech recognition. By randomly masking bands in the log Mel spectogram this method leads to impressive performance improvements. In this paper, we investigate the usage of SpecAugment for speaker verification tasks. Two different models,…

Cited by 0SourceScholar
2018

Analysis of Multilingual Blstm Acoustic Model on Low and High Resource Languages

ICASSP 2018accepted

The paper provides an analysis of automatic speech recognition systems (ASR) based on multilingual BLSTM, where we used multi-task training with separate classification layer for each language. The focus is on low resource languages, where only a limited amount of transcribed speech is available. In…

Cited by 0SourceScholar
2018

Optimization of Speaker-Aware Multichannel Speech Extraction with ASR Criterion

ICASSP 2018accepted

This paper addresses the problem of recognizing speech corrupted by overlapping speakers in a multichannel setting. To extract a target speaker from the mixture, we use a neural network based beamformer which uses masks estimated by a neural network to compute statistically optimal spatial filters.…

Cited by 0SourceScholar
2017

Bayesian phonotactic Language Model for Acoustic Unit Discovery

ICASSP 2017accepted

Recent work on Acoustic Unit Discovery (AUD) has led to the development of a non-parametric Bayesian phone-loop model where the prior over the probability of the phone-like units is assumed to be sampled from a Dirichlet Process (DP). In this work, we propose to improve this model by incorporating a…

Cited by 0SourceScholar
2017

Residual memory networks: Feed-forward approach to learn long-term temporal dependencies

ICASSP 2017accepted

Training deep recurrent neural network (RNN) architectures is complicated due to the increased network complexity. This disrupts the learning of higher order abstracts using deep RNN. In case of feed-forward networks training deep structures is simple and faster while learning long-term temporal inf…

Cited by 0SourceScholar
2017

Topic identification of spoken documents using unsupervised acoustic unit discovery

ICASSP 2017accepted

This paper investigates the application of unsupervised acoustic unit discovery for topic identification (topic ID) of spoken audio documents. The acoustic unit discovery method is based on a non-parametric Bayesian phone-loop model that segments a speech utterance into phone-like categories. The di…

Cited by 0SourceScholar
2016

Analysis of DNN approaches to speaker identification

ICASSP 2016accepted

This work studies the usage of the Deep Neural Network (DNN) Bottleneck (BN) features together with the traditional MFCC features in the task of i-vector-based speaker recognition. We decouple the sufficient statistics extraction by using separate GMM models for frame alignment, and for statistics n…

Cited by 0SourceScholar
2015

Copingwith channel mismatch in Query-by-Example - But QUESST 2014

ICASSP 2015accepted

The paper investigates into Query by Example (QbE) - a spoken term detection technique with queries entered by voice. It describes BUT QbE system that achieved the best accuracy in MediaEval QUESST2014 evaluations. This evaluation was challenging because of severe mismatch between queries and uttera…

Cited by 0SourceScholar