← Search

Martin Kocour

5 accepted papers

2026

ADAPTING DIARIZATION-CONDITIONED WHISPER FOR END-TO-END MULTI-TALKER SPEECH RECOGNITION

ICASSP 2026oral

We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned Whisper (DiCoW) encoder to extract target-speaker embeddings, which are concatena…

Cited by 0SourcePDFScholar
2025

Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition

ICASSP 2025accepted

This paper presents an efficient decoding approach for end-to-end automatic speech recognition (E2E-ASR) with large language models (LLMs). Although shallow fusion is the most common approach to incorporate language models into E2E-ASR decoding, we face two practical problems with LLMs. (1) LLM infe…

Cited by 0SourceScholar
2022

Call-Sign Recognition and Understanding for Noisy Air-Traffic Transcripts Using Surveillance Information

ICASSP 2022accepted

Air traffic control (ATC) relies on communication via speech between pilot and air-traffic controller (ATCO). The call-sign, as unique identifier for each flight, is used to address a specific pilot by the ATCO. Extracting the call-sign from the communication is a challenge because of the noisy ATC…

Cited by 0SourceScholar
2022

GPU-Accelerated Forward-Backward Algorithm with Application to Lattice-Free MMI

ICASSP 2022accepted

We propose to express the forward-backward algorithm in terms of operations between sparse matrices in a specific semiring. This new perspective naturally leads to a GPU-friendly algorithm which is easy to implement in Julia or any programming languages with native support of semiring algebra. We us…

Cited by 0SourceScholar