← Search

Alexander Polok

3 accepted papers

2026

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

ICASSP 2026oral

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on specific domains, few systems generalize well across out-of-domain datasets. Our prior work, Diarization-Conditioned Whis…

Cited by 0SourcePDFScholar
2025

Target Speaker ASR with Whisper

ICASSP 2025accepted

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs than to learn th…

Cited by 0SourceScholar