← Search

Konrad Kowalczyk

8 accepted papers

2025

Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio

ICASSP 2025accepted

Hallucinations of deep neural models are amongst key challenges in automatic speech recognition (ASR). In this paper, we investigate hallucinations of the Whisper ASR model induced by non-speech audio segments present during inference. By inducting hallucinations with various types of sounds, we sho…

Cited by 0SourceScholar
2022

Convolutional Weighted Minimum Mean Square Error Filter for Joint Source Separation and Dereverberation

ICASSP 2022accepted

Practical scenarios with multiple simultaneously active speakers recorded using one or more microphones in reverberant rooms pose a challenging problem when the extraction of the desired speaker signal is sought for. The majority of techniques found in the literature facilitate either source separat…

Cited by 0SourceScholar
2022

Wishart Localization Prior On Spatial Covariance Matrix In Ambisonic Source Separation Using Non-Negative Tensor Factorization

ICASSP 2022accepted

This paper presents an extension of the existing Non-negative Tensor Factorization (NTF) based method for sound source separation under reverberant conditions, formulated for Ambisonic microphone mixture signals. In particular, we address the problem of optimal exploitation of the prior knowledge co…

Cited by 0SourceScholar
2021

Maximum a Posteriori Estimator for Convolutive Sound Source Separation with Sub-Source Based NTF Model and the Localization Probabilistic Prior on the Mixing Matrix

ICASSP 2021accepted

In this paper we present a method for the separation of sound source signals recorded using multiple microphones in a reverberant room. In particular, we propose a maximum a posteriori (MAP) estimator based on the multichannel nonnegative tensor factorization (NTF) model with the localization prior…

Cited by 0SourceScholar
2015

Residual noise control using a parametric multichannel Wiener filter

ICASSP 2015accepted

Multichannel noise reduction techniques are commonly used in speech communication applications. In these applications, it is often desired to maintain a residual amount of background noise to avoid perceptually unpleasant artifacts, such as musical tones or time periods of complete silence. Noise re…

Cited by 0SourceScholar