← Search

Hynek Hermansky

9 accepted papers

2020

A Practical Two-Stage Training Strategy for Multi-Stream End-to-End Speech Recognition

ICASSP 2020accepted

The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study offered a promising direction within end-to-end automatic speech recognition, where parallel encoders aim to capture dive…

Cited by 0SourceScholar
2019

Deriving Spectro-temporal Properties of Hearing from Speech Data

ICASSP 2019accepted

Human hearing and human speech are intrinsically tied together, as the properties of speech almost certainly developed in order to be heard by human ears. As a result of this connection, it has been shown that certain properties of human hearing are mimicked within data-driven systems that are train…

Cited by 0SourceScholar
2019

M-vectors: Sub-band Based Energy Modulation Features for Multi-stream Automatic Speech Recognition

ICASSP 2019accepted

In this paper, we propose a novel method to capture energy modulations from different frequency bands in speech into frame-level feature vectors, called Modulation-vectors or M-vectors, for use in Automatic Speech Recognition (ASR) systems. We show that in different multi-stream setups, with paralle…

Cited by 0SourceScholar
2019

Stream Attention-based Multi-array End-to-end Speech Recognition

ICASSP 2019accepted

Automatic Speech Recognition (ASR) using multiple microphone arrays has achieved great success in the far-field robustness. Taking advantage of all the information that each array shares and contributes is crucial in this task. Motivated by the advances of joint Connectionist Temporal Classification…

Cited by 21SourceScholar
2019

Towards Automatic Methods to Detect Errors in Transcriptions of Speech Recordings

ICASSP 2019accepted

This work explores different methods to detect errors in transcriptions of speech recordings. We artificially corrupt well transcribed speech transcriptions with three types of errors: substitution, insertion and deletion on TIMIT phonemic transcriptions and WSJ word transcriptions. First, we use Ba…

Cited by 0SourceScholar
2017

Predicting error rates for unknown data in automatic speech recognition

ICASSP 2017accepted

In this paper we investigate methods to predict word error rates in automatic speech recognition in the presence of unknown noise types, which have not been seen during training. The performance measures operate on phoneme posteriorgrams that are obtained from neural nets. We compare average frame-w…

Cited by 0SourceScholar
2015

Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial Workshop

ICASSP 2015accepted

A group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which…

Cited by 0SourceScholar