← Search

Petr Motlícek

21 accepted papers

2025

Speech Data Selection for Efficient ASR Fine-Tuning using Domain Classifier and Pseudo-Label Filtering

ICASSP 2025accepted

In real-world speech data processing, the scarcity of annotated data and the abundance of unlabelled speech data present a significant challenge. To address this, we propose an efficient data selection pipeline for fine-tuning ASR models by generating pseudo-labels using WhisperX pipeline and select…

Cited by 6SourceScholar
2025

XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models

ICASSP 2025accepted

Self-supervised pretrained models exhibit competitive performance in automatic speech recognition (ASR) on finetuning, even with limited in-domain supervised data. However, popular pretrained models are not suitable for streaming ASR because they are trained with full attention context. In this pape…

Cited by 0SourceScholar
2024

Contextual Biasing Methods for Improving Rare Word Detection in Automatic Speech Recognition

ICASSP 2024accepted

In specialized domains like Air Traffic Control (ATC), a notable challenge in porting a deployed Automatic Speech Recognition (ASR) system from one airport to another is the alteration in the set of crucial words that must be accurately detected in the new environment. Typically, such words have lim…

Cited by 0SourceScholar
2024

Fine-Tuning Self-Supervised Models for Language Identification Using Orthonormal Constraint

ICASSP 2024accepted

Self-supervised models trained with high linguistic diversity, such as the XLS-R model, can be effectively fine-tuned for the language recognition task. Typically, a back-end classifier followed by statistics pooling layer are added during training. Commonly used back-end classifiers require a large…

Cited by 7SourceScholar
2024

Multitask Speech Recognition and Speaker Change Detection for Unknown Number of Speakers

ICASSP 2024accepted

Traditionally, automatic speech recognition (ASR) and speaker change detection (SCD) systems have been independently trained to generate comprehensive transcripts accompanied by speaker turns. Recently, joint training of ASR and SCD systems, by inserting speaker turn tokens in the ASR training text,…

Cited by 0SourceScholar
2024

Probability-Aware Word-Confusion-Network-To-Text Alignment Approach for Intent Classification

ICASSP 2024accepted

Spoken Language Understanding (SLU) technologies have greatly improved due to the effective pretraining of speech representations. A common requirement of industry-based solutions is the portability to deploy SLU models in voice-assistant devices. Thus, distilling knowledge from large text-based lan…

Cited by 1SourceScholar
2023

Effectiveness of Text, Acoustic, and Lattice-Based Representations in Spoken Language Understanding Tasks

ICASSP 2023accepted

In this paper, we perform an exhaustive evaluation of different representations to address the intent classification problem in a Spoken Language Understanding (SLU) setup. We benchmark three types of systems to perform the SLU intent detection task: 1) text-based, 2) lattice-based, and a novel 3) m…

Cited by 0SourceScholar
2022

A Two-Step Approach to Leverage Contextual Data: Speech Recognition in Air-Traffic Communications

ICASSP 2022accepted

Automatic Speech Recognition (ASR), as the assistance of speech communication between pilots and air-traffic controllers, can significantly reduce the complexity of the task and increase the reliability of transmitted information. ASR application can lead to a lower number of incidents caused by mis…

Cited by 0SourceScholar
2021

A Comparison of Methods for OOV-Word Recognition on a New Public Dataset

ICASSP 2021accepted

A common problem for automatic speech recognition systems is how to recognize words that they did not see during training. Currently there is no established method of evaluating different techniques for tackling this problem.We propose using the CommonVoice dataset to create test sets for multiple l…

Cited by 0SourceScholar
2020

Incremental Semi-Supervised Learning for Multi-Genre Speech Recognition

ICASSP 2020accepted

In this work, we explore a data scheduling strategy for semi-supervised learning (SSL) for acoustic modeling in automatic speech recognition. The conventional approach uses a seed model trained with supervised data to automatically recognize the entire set of unlabeled (auxiliary) data to generate n…

Cited by 0SourceScholar
2019

A Bayesian Approach to Inter-task Fusion for Speaker Recognition

ICASSP 2019accepted

In i-vector based speaker recognition systems, back-end classifiers are trained to factor out nuisance information and retain only the speaker identity. As a result, variabilities arising due to gender, language and accent (among many others) are suppressed. Inter-task fusion, in which such metadata…

Cited by 0SourceScholar
2019

Adaptation of Multiple Sound Source Localization Neural Networks with Weak Supervision and Domain-adversarial Training

ICASSP 2019accepted

Despite the recent success of deep neural network-based approaches in sound source localization, these approaches suffer the limitations that the required annotation process is costly, and the mismatch between the training and test conditions undermines the performance. This paper addresses the ques…

Cited by 0SourceScholar
2018

DNN Based Speaker Embedding Using Content Information for Text-Dependent Speaker Verification

ICASSP 2018accepted

In this paper, we are interested in exploring Deep Neural Network (DNN) based speaker embedding for Random-digit task using content information. To this end, a technique is applied to automatically select common phonetic units between the enrollment and test data to produce speaker verification scor…

Cited by 0SourceScholar
2017

Exploiting sequence information for text-dependent Speaker Verification

ICASSP 2017accepted

Model-based approaches to Speaker Verification (SV), such as Joint Factor Analysis (JFA), i-vector and relevance Maximum-a-Posteriori (MAP), have shown to provide state-of-the-art performance for text-dependent systems with fixed phrases. The performance of i-vector and JFA models has been further e…

Cited by 0SourceScholar
2017

Intra-class covariance adaptation in PLDA back-ends for speaker verification

ICASSP 2017accepted

Multi-session training conditions are becoming increasingly common in recent benchmark datasets for both text-independent and text-dependent speaker verification. In the state-of-the-art i-vector framework for speaker verification, such conditions are addressed by simple techniques such as averaging…

Cited by 0SourceScholar
2016

Deep neural network based posteriors for text-dependent speaker verification

ICASSP 2016accepted

The i-vector and Joint Factor Analysis (JFA) systems for text-dependent speaker verification use sufficient statistics computed from a speech utterance to estimate speaker models. These statistics average the acoustic information over the utterance thereby losing all the sequence information. In thi…

Cited by 0SourceScholar
2016

Information theoretic clustering for unsupervised domain-adaptation

ICASSP 2016accepted

The aim of the domain-adaptation task for speaker verification is to exploit unlabelled target domain data by using the labelled source domain data effectively. The i-vector based Probabilistic Linear Discriminant Analysis (PLDA) framework approaches this task by clustering the target domain data an…

Cited by 0SourceScholar
2016

System fusion and speaker linking for longitudinal diarization of TV shows

ICASSP 2016accepted

Performing speaker diarization while uniquely identifying the speakers in a collection of audio recordings is a challenging task. Based on our previous work on speaker diarization and linking, we developed a system for diarizing longitudinal TV show data sets based on the fusion of speaker diarizati…

Cited by 0SourceScholar
2015

Combining SGMM speaker vectors and KL-HMM approach for speaker diarization

ICASSP 2015accepted

In this paper, a method to use SGMM speaker vectors for speaker diarization is introduced. The architecture of the Information Bottleneck (IB) based speaker diarization is utilized for this purpose. The audio for speaker diarization is split into short uniform segments. Speaker vectors are obtained…

Cited by 5SourceScholar
2015

Employment of Subspace Gaussian Mixture Models in speaker recognition

ICASSP 2015accepted

This paper presents Subspace Gaussian Mixture Model (SGMM) approach employed as a probabilistic generative model to estimate speaker vector representations to be subsequently used in the speaker verification task. SGMMs have already been shown to significantly outperform traditional HMM/GMMs in Auto…

Cited by 29SourceScholar
2015

Learning feature mapping using deep neural network bottleneck features for distant large vocabulary speech recognition

ICASSP 2015accepted

Automatic speech recognition from distant microphones is a difficult task because recordings are affected by reverberation and background noise. First, the application of the deep neural network (DNN)/hidden Markov model (HMM) hybrid acoustic models for distant speech recognition task using AMI meet…

Cited by 0SourceScholar