← Search

Bernd T. Meyer

9 accepted papers

2023

Multilingual Query-by-Example Keyword Spotting with Metric Learning and Phoneme-to-Embedding Mapping

ICASSP 2023accepted

In this paper, we propose a multilingual query-by-example keyword spotting (KWS) system based on a residual neural network. The model is trained as a classifier on a multilingual keyword dataset extracted from Common Voice sentences and fine-tuned using circle loss. We demonstrate the generalization…

Cited by 0SourceScholar
2021

Non-Intrusive Binaural Prediction of Speech Intelligibility Based on Phoneme Classification

ICASSP 2021accepted

In this study, we explore an approach for modeling speech intelligibility in spatial acoustic scenes. To this end, we combine a non-intrusive binaural frontend with a deep neural network (DNN) borrowed from a standard automatic speech recognition (ASR) system. The DNN estimates phoneme probabilities…

Cited by 0SourceScholar
2020

DNN-Based Speech Presence Probability Estimation for Multi-Frame Single-Microphone Speech Enhancement

ICASSP 2020accepted

Multi-frame approaches for single-microphone speech enhancement, e.g., the multi-frame minimum-power-distortionless-response (MFMPDR) filter, are able to exploit speech correlations across neighboring time frames. In contrast to single-frame approaches such as the Wiener gain, it has been shown that…

Cited by 0SourceScholar
2019

Improving Deep Models of Speech Quality Prediction through Voice Activity Detection and Entropy-based Measures

ICASSP 2019accepted

This paper explores Deep machine listening for Estimating Speech Quality (DESQ), which predicts the perceived speech quality based on phoneme posterior probabilities obtained from a deep neural network. The degradation of phonemes is quantified with the entropy-based Gini measure that is compared to…

Cited by 0SourceScholar
2017

Combination strategy based on relative performance monitoring for multi-stream reverberant speech recognition

ICASSP 2017accepted

A multi-stream framework with deep neural network (DNN) classifiers is applied to improve automatic speech recognition (ASR) in environments with different reverberation characteristics. We propose a room parameter estimation model to establish a reliable combination strategy which performs on eithe…

Cited by 0SourceScholar
2017

On DNN posterior probability combination in multi-stream speech recognition for reverberant environments

ICASSP 2017accepted

A multi-stream framework with deep neural network (DNN) classifiers has been applied in this paper to improve automatic speech recognition (ASR) performance in environments with different reverberation characteristics. We propose a room parameter estimation model to determine the stream weights for…

Cited by 0SourceScholar
2017

Predicting error rates for unknown data in automatic speech recognition

ICASSP 2017accepted

In this paper we investigate methods to predict word error rates in automatic speech recognition in the presence of unknown noise types, which have not been seen during training. The performance measures operate on phoneme posteriorgrams that are obtained from neural nets. We compare average frame-w…

Cited by 0SourceScholar
2015

A study on joint beamforming and spectral enhancement for robust speech recognition in reverberant environments

ICASSP 2015accepted

This work evaluates multi-microphone beamforming and single-microphone spectral enhancement strategies to alleviate the reverberation effect for robust automatic speech recognition (ASR) systems in different reverberant environments characterized by different reverberation times T60 and direct-to-re…

Cited by 0SourceScholar