← Search

Jan Trmal

5 accepted papers

2024

ConEC: Earnings Call Dataset with Real-world Contexts for Benchmarking Contextual Speech Recognition

COLING 2024main

Knowing the particular context associated with a conversation can help improving the performance of an automatic speech recognition (ASR) system. For example, if we are provided with a list of in-context words or phrases — such as the speaker’s contacts or recent song playlists — during inference, w…

2023

Building Keyword Search System from End-To-End Asr Systems

ICASSP 2023accepted

Keyword search (KWS) systems are commonly built on top of existing automatic speech recognition (ASR) systems. However, end-to-end (E2E) ASR models are not naturally equipped with word-level timing information or confidence. Existing methods for re-purposing E2E ASR systems for KWS are largely heuri…

Cited by 0SourceScholar
2020

Multi-Task Self-Supervised Learning for Robust Speech Recognition

ICASSP 2020accepted

Despite the growing interest in unsupervised learning, extracting meaningful knowledge from unlabelled audio remains an open challenge. To take a step in this direction, we recently proposed a problem-agnostic speech encoder (PASE), that combines a convolutional encoder followed by multiple neural n…

Cited by 0SourceScholar
2018

On the Use of Grapheme Models for Searching in Large Spoken Archives

ICASSP 2018accepted

This paper explores the possibility to use grapheme-based word and sub-word models in the task of spoken term detection (STD). The usage of grapheme models eliminates the need for expert-prepared pronunciation lexicons (which are often far from complete) and/or trainable grapheme-to-phoneme (G2P) al…

Cited by 0SourceScholar
2017

Bayesian joint-sequence models for grapheme-to-phoneme conversion

ICASSP 2017accepted

We describe a fully Bayesian approach to grapheme-to-phoneme conversion based on the joint-sequence model (JSM). Usually, standard smoothed n-gram language models (LM, e.g. Kneser-Ney) are used with JSMs to model graphone sequences (joint grapheme-phoneme pairs). However, we take a Bayesian approach…

Cited by 0SourceScholar