← Search

Jan Honza Cernocký

9 accepted papers

2022

Multi-Channel Speaker Verification with Conv-Tasnet Based Beamformer

ICASSP 2022accepted

We focus on the problem of speaker recognition in far-field multichannel data. The main contribution is introducing an alternative way of predicting spatial covariance matrices (SCMs) for a beamformer from the time domain signal. We propose to use ConvTasNet, a well-known source separation model, an…

Cited by 0SourceScholar
2022

Multisv: Dataset for Far-Field Multi-Channel Speaker Verification

ICASSP 2022accepted

Motivated by unconsolidated data situation and the lack of a standard benchmark in the field, we complement our previous efforts and present a comprehensive corpus designed for training and evaluating text-independent multi-channel speaker verification systems. It can be readily used also for experi…

Cited by 0SourceScholar
2021

Analysis of X-Vectors for Low-Resource Speech Recognition

ICASSP 2021accepted

The paper presents a study of usability of x-vectors for adaptation of automatic speech recognition (ASR) systems. X-vectors are Neural Network (NN)-based speaker embeddings recently proposed in speaker recognition (SR). They quickly replaced common i-vectors and became new state-of-the-art techniqu…

Cited by 0SourceScholar
2021

Eat: Enhanced ASR-TTS for Self-Supervised Speech Recognition

ICASSP 2021accepted

Self-supervised ASR-TTS models suffer in out-of-domain data conditions. Here we propose an enhanced ASR-TTS (EAT) model that incorporates two main features: 1) The ASR→TTS direction is equipped with a language model reward to penalize the ASR hypotheses before forwarding it to TTS. 2) In the TTS→ASR…

Cited by 0SourceScholar
2019

How to Improve Your Speaker Embeddings Extractor in Generic Toolkits

ICASSP 2019accepted

Recently, speaker embeddings extracted with deep neural networks became the state-of-the-art method for speaker verification. In this paper we aim to facilitate its implementation on a more generic toolkit than Kaldi, which we anticipate to enable further improvements on the method. We examine sever…

Cited by 51SourceScholar
2019

Promising Accurate Prefix Boosting for Sequence-to-sequence ASR

ICASSP 2019accepted

In this paper, we present promising accurate prefix boosting (PAPB), a discriminative training technique for attention based sequence-to-sequence (seq2seq) ASR. PAPB is devised to unify the training and testing scheme effectively. The training procedure involves maximizing the score of each partial…

Cited by 16SourceScholar
2018

Dereverberation and Beamforming in Far-Field Speaker Recognition

ICASSP 2018accepted

This paper deals with far-field speaker recognition. On a corpus of NIST SRE 2010 data retransmitted in a real room with multiple microphones, we first demonstrate how room acoustics cause significant degradation of state-of-the-art i-vector based speaker recognition system. We then investigate seve…

Cited by 0SourceScholar
2016

Multilingual region-dependent transforms

ICASSP 2016accepted

In recent years, trained feature extraction (FE) schemes based on neural networks have replaced or complemented traditional approaches in top performing systems. This paper deals with FE in multilingual scenarios with a target language with low amount of transcribed data. Continuing our previous wor…

Cited by 0SourceScholar
2016

Sequence summarizing neural network for speaker adaptation

ICASSP 2016accepted

In this paper, we propose a DNN adaptation technique, where the i-vector extractor is replaced by a Sequence Summarizing Neural Network (SSNN). Similarly to i-vector extractor, the SSNN produces a "summary vector", representing an acoustic summary of an utterance. Such vector is then appended to the…

Cited by 0SourceScholar