← Search

Ondrej Klejch

9 accepted papers

2026

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

ICLR 2026oral

Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but rarely validated against subjective ones. Both kinds of metrics are challenged by…

Cited by 0SourcecodeScholar
2025

Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives

ICASSP 2025accepted

We develop automatic speech recognition (ASR) systems for stories told by Afrikaans and isiXhosa preschool children. Oral narratives provide a way to assess children’s language development before they learn to read. We consider a range of prior child-speech ASR strategies to determine which is best…

Cited by 0SourceScholar
2025

Spoken Document Retrieval for an Unwritten Language: A Case Study on Gormati

EMNLP 2025

Speakers of unwritten languages have the potential to benefit from speech-based automatic information retrieval systems. This paper proposes a speech embedding technique that facilitates such a system that we can be used in a zero-shot manner on the target language. After conducting development expe

Cited by 0SourcePDFScholar
2024

Speech Collage: Code-Switched Audio Generation by Collaging Monolingual Corpora

ICASSP 2024accepted

Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a method that synthesizes CS data from monolingual corpora by splicing audio segme…

Cited by 0SourceScholar
2023

Efficient Intelligibility Evaluation Using Keyword Spotting: A Study on Audio-Visual Speech Enhancement

ICASSP 2023accepted

We propose a new method for human speech intelligibility evaluation based on keyword spotting. In this method, participants play a stimulus and select the word they hear from a close set of alternatives. To find which sentence to use, the target word, and alternatives we mine a large set of stimuli…

Cited by 0SourceScholar
2023

The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR

ICASSP 2023accepted

English is the most widely spoken language in the world, used daily by millions of people as a first or second language in many different contexts. As a result, there are many varieties of English. Although the great many advances in English automatic speech recognition (ASR) over the past decades,…

Cited by 0SourceScholar
2023

Towards Zero-Shot Code-Switched Speech Recognition

ICASSP 2023accepted

In this work, we seek to build effective code-switched (CS) automatic speech recognition systems (ASR) under the zero-shot set-ting where no transcribed CS speech data is available for training. Previously proposed frameworks which conditionally factorize the bilingual task into its constituent mono…

Cited by 0SourceScholar
2020

Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection

ICASSP 2020accepted

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large, carefully labeled audio-visual active speaker dataset has limited ev…

Cited by 0SourceScholar
2017

Sequence-to-sequence models for punctuated transcription combining lexical and acoustic features

ICASSP 2017accepted

In this paper we present an extension of our previously described neural machine translation based system for punctuated transcription. This extension allows the system to map from per frame acoustic features to word level representations by replacing the traditional encoder in the encoder-decoder a…

Cited by 0SourceScholar