← Search

Martin Karafiát

8 accepted papers

2026

ADAPTING DIARIZATION-CONDITIONED WHISPER FOR END-TO-END MULTI-TALKER SPEECH RECOGNITION

ICASSP 2026oral

We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned Whisper (DiCoW) encoder to extract target-speaker embeddings, which are concatena…

Cited by 2SourcePDFScholar
2021

Analysis of X-Vectors for Low-Resource Speech Recognition

ICASSP 2021accepted

The paper presents a study of usability of x-vectors for adaptation of automatic speech recognition (ASR) systems. X-vectors are Neural Network (NN)-based speaker embeddings recently proposed in speaker recognition (SR). They quickly replaced common i-vectors and became new state-of-the-art techniqu…

Cited by 0SourceScholar
2021

Jointly Trained Transformers Models for Spoken Language Translation

ICASSP 2021accepted

End-to-End and cascade (ASR-MT) spoken language translation (SLT) systems are reaching comparable performances, however, a large degradation is observed when translating the ASR hypothesis in comparison to using oracle input text. In this work, degradation in performance is reduced by creating an En…

Cited by 0SourceScholar
2019

Promising Accurate Prefix Boosting for Sequence-to-sequence ASR

ICASSP 2019accepted

In this paper, we present promising accurate prefix boosting (PAPB), a discriminative training technique for attention based sequence-to-sequence (seq2seq) ASR. PAPB is devised to unify the training and testing scheme effectively. The training procedure involves maximizing the score of each partial…

Cited by 16SourceScholar
2018

Analysis of Multilingual Blstm Acoustic Model on Low and High Resource Languages

ICASSP 2018accepted

The paper provides an analysis of automatic speech recognition systems (ASR) based on multilingual BLSTM, where we used multi-task training with separate classification layer for each language. The focus is on low resource languages, where only a limited amount of transcribed speech is available. In…

Cited by 0SourceScholar
2017

Residual memory networks: Feed-forward approach to learn long-term temporal dependencies

ICASSP 2017accepted

Training deep recurrent neural network (RNN) architectures is complicated due to the increased network complexity. This disrupts the learning of higher order abstracts using deep RNN. In case of feed-forward networks training deep structures is simple and faster while learning long-term temporal inf…

Cited by 0SourceScholar
2016

Multilingual region-dependent transforms

ICASSP 2016accepted

In recent years, trained feature extraction (FE) schemes based on neural networks have replaced or complemented traditional approaches in top performing systems. This paper deals with FE in multilingual scenarios with a target language with low amount of transcribed data. Continuing our previous wor…

Cited by 0SourceScholar
2016

Sequence summarizing neural network for speaker adaptation

ICASSP 2016accepted

In this paper, we propose a DNN adaptation technique, where the i-vector extractor is replaced by a Sequence Summarizing Neural Network (SSNN). Similarly to i-vector extractor, the SSNN produces a "summary vector", representing an acoustic summary of an utterance. Such vector is then appended to the…

Cited by 0SourceScholar