← Search

Stefan Goetze

16 accepted papers

2025

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings

ICASSP 2025accepted

Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the synchronisation between speech and lip movement. Recent work h…

Cited by 0SourceScholar
2024

Active Learning for Sound Event Classification Using Bayesian Neural Networks with Gaussian Variational Posterior

ICASSP 2024accepted

Manual annotation of audio material is cumbersome. Active learning aims at minimizing the annotation effort by iteratively selecting an acquisition batch of unlabeled data, asking a human to annotate the selected data and re-training a classifier until an annotation budget is depleted. In this paper…

Cited by 3SourceScholar
2024

Combining Conformer and Dual-Path-Transformer Networks for Single Channel Noisy Reverberant Speech Separation

ICASSP 2024accepted

Separation of overlapping speakers remains an active area of speech technology research. Many deep neural network (DNN) separation models propose modelling local and global temporal context separately using alternating DNN layers. Two such models are SepFormer and TD-Conformer. The largest configura…

Cited by 0SourceScholar
2024

Multi-CMGAN+/+: Leveraging Multi-Objective Speech Quality Metric Prediction for Speech Enhancement

ICASSP 2024accepted

Neural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on artificially created labelled training data such that the ne…

Cited by 0SourceScholar
2024

Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory Models

ICASSP 2024accepted

Neural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work…

Cited by 0SourceScholar
2024

Refining Text Input For Augmentative and Alternative Communication (AAC) Devices: Analysing Language Model Layers For Optimisation

ICASSP 2024accepted

Communication impairments are prevalent among a significant proportion of individuals. Methods of Augmentative and Alternative Communication (AAC) can support people with speech disorders (PwSD) to some extent, but AAC users encounter substantial difficulties when engaging in open-domain social inte…

Cited by 0SourceScholar
2023

Deformable Temporal Convolutional Networks for Monaural Noisy Reverberant Speech Separation

ICASSP 2023accepted

Speech separation models are used for isolating individual speakers in many speech processing applications. Deep learning models have been shown to lead to state-of-the-art (SOTA) results on a number of speech separation benchmarks. One such class of models known as temporal convolutional networks (…

Cited by 0SourceScholar
2023

Moving Towards Non-Binary Gender Identification Via Analysis of System Errors in Binary Gender Classification

ICASSP 2023accepted

This paper aims to analyse human perceptions of gender in speech signals, focusing on signals that are misclassified by methods for binary gender classification, looking at the features of speech signals that are more likely to be misclassified, or classified as either nonbinary or unclassifiable. T…

Cited by 0SourceScholar
2023

Perceive and Predict: Self-Supervised Speech Representation Based Loss Functions for Speech Enhancement

ICASSP 2023accepted

Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the deepest or final outputs of self supervised speech representation models, rather…

Cited by 0SourceScholar
2017

Combination strategy based on relative performance monitoring for multi-stream reverberant speech recognition

ICASSP 2017accepted

A multi-stream framework with deep neural network (DNN) classifiers is applied to improve automatic speech recognition (ASR) in environments with different reverberation characteristics. We propose a room parameter estimation model to establish a reliable combination strategy which performs on eithe…

Cited by 0SourceScholar
2017

Measuring, modelling and predicting perceived reverberation

ICASSP 2017accepted

This paper investigates the relationship between the perceived level of reverberation and parameters measured from the room impulse response (RIR), as well as the design of an instrumental measure that predicts this perceived level. We first present the results of an experimental listening test cond…

Cited by 0SourceScholar
2017

On DNN posterior probability combination in multi-stream speech recognition for reverberant environments

ICASSP 2017accepted

A multi-stream framework with deep neural network (DNN) classifiers has been applied in this paper to improve automatic speech recognition (ASR) performance in environments with different reverberation characteristics. We propose a room parameter estimation model to determine the stream weights for…

Cited by 0SourceScholar
2016

Classification of human cough signals using spectro-temporal Gabor filterbank features

ICASSP 2016accepted

This contribution investigates the use of features derived from a Gabor filterbank (GFB) for the application of acoustic cough classification. Gabor filters are two-dimensional filters that decompose the spectro-temporal power density further into components which capture spectral, temporal and join…

Cited by 0SourceScholar
2016

Perceptual and instrumental evaluation of the perceived level of reverberation

ICASSP 2016accepted

Perceptual measures are usually considered more reliable than instrumental measures for evaluating the perceived level of reverberation. However, such measures are costly in both time and money, and, due to variations in stimuli or assessors, the resulting data is not always statistically significan…

Cited by 13SourceScholar
2015

A study on joint beamforming and spectral enhancement for robust speech recognition in reverberant environments

ICASSP 2015accepted

This work evaluates multi-microphone beamforming and single-microphone spectral enhancement strategies to alleviate the reverberation effect for robust automatic speech recognition (ASR) systems in different reverberant environments characterized by different reverberation times T60 and direct-to-re…

Cited by 3SourceScholar