← Search

Mike Brookes

16 accepted papers

2024

Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks

ICASSP 2024accepted

Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an enco…

Cited by 0SourceScholar
2024

Speech Enhancement in Hearing Aids Using Target Speech Presence Estimation Based on a Delayed Remote Microphone Signal

ICASSP 2024accepted

Speech enhancement in hearing aids (HAs) can take advantage of a wireless remote microphone (RM) having a better signal-to-noise ratio than the HA microphones. However, using the RM effectively is complicated by the time delay between the acoustic and wireless signals. Methods in the literature assu…

Cited by 0SourceScholar
2023

Graph Neural Networks for Sound Source Localization on Distributed Microphone Networks

ICASSP 2023accepted

Distributed Microphone Arrays (DMAs) present many challenges with respect to centralized microphone arrays. An important requirement of applications on these arrays is handling a variable number of input channels. We consider the use of Graph Neural Networks (GNNs) as a solution to this challenge. W…

Cited by 0SourceScholar
2023

The MBSTOI Binaural Intelligibility Metric Using a Close-Talking Microphone Reference

ICASSP 2023accepted

Intelligibility metrics are a fast way to determine how comprehensible a target signal is in a noisy situation. Most metrics however rely on having a clean reference signal for computation and are not adapted to live recordings. In this paper the deep correlation modified binaural short time objecti…

Cited by 0SourceScholar
2021

Processing Pipelines for Efficient, Physically-Accurate Simulation of Microphone Array Signals in Dynamic Sound Scenes

ICASSP 2021accepted

Multichannel acoustic signal processing is predicated on the fact that the interchannel relationships between the received signals can be exploited to infer information about the acoustic scene. Recently there has been increasing interest in algorithms which are applicable in dynamic scenes, where t…

Cited by 2SourceScholar
2018

Room Identification Using Frequency Dependence of Spectral Decay Statistics

ICASSP 2018accepted

A method for room identification is proposed based on the reverberation properties of multichannel speech recordings. The approach exploits the dependence of spectral decay statistics on the reverberation time of a room. The average negative-side variance within 1/3-octave bands is proposed as the i…

Cited by 4SourceScholar
2017

Frequency-domain under-modelled blind system identification based on cross power spectrum and sparsity regularization

ICASSP 2017accepted

In room acoustics, under-modelled multichannel blind system identification (BSI) aims to estimate the early part of the room impulse responses (RIRs), and it can be widely used in applications such as speaker localization, room geometry identification and beamforming based speech dereverberation. In…

Cited by 0SourceScholar
2017

Identifying a multiple plane plenoptic function from a swiped image

ICASSP 2017accepted

Blur in images, caused by camera motion with an open shutter, is usually thought of as a problem. The algorithm described in this paper shows instead that it is possible to use the blur caused by the integration of light rays at different locations along a moving camera trajectory to extract informa…

Cited by 0SourceScholar
2017

Improving the perceptual quality of ideal binary masked speech

ICASSP 2017accepted

It is known that applying a time-frequency binary mask to very noisy speech can improve its intelligibility but results in poor perceptual quality. In this paper we propose a new approach to applying a binary mask that combines the intelligibility gains of conventional binary masking with the percep…

Cited by 0SourceScholar
2017

Robust spherical harmonic domain interpolation of spatially sampled array manifolds

ICASSP 2017accepted

Accurate interpolation of the array manifold is an important first step for the acoustic simulation of rapidly moving microphone arrays. Spherical harmonic domain interpolation has been proposed and well studied in the context of head-related transfer functions but has focussed on perceptual, rather…

Cited by 0SourceScholar
2016

Speech enhancement using an MMSE spectral amplitude estimator based on a modulation domain Kalman filter with a Gamma prior

ICASSP 2016accepted

In this paper, we propose a minimum mean square error spectral estimator for clean speech spectral amplitudes that uses a Kalman filter to model the temporal dynamics of the spectral amplitudes in the modulation domain. Using a two-parameter Gamma distribution to model the prior distribution of the…

Cited by 0SourceScholar
2015

SOBM - a binary mask for noisy speech that optimises an objective intelligibility metric

ICASSP 2015accepted

It is known that the intelligibility of noisy speech can be improved by applying a binary-valued gain mask to a time-frequency representation of the speech. We present the SOBM, an oracle binary mask that maximises STOI, an objective speech intelligibility metric. We show how to determine the SOBM f…

Cited by 0SourceScholar
2015

Single-channel blind estimation of reverberation parameters

ICASSP 2015accepted

The reverberation of an acoustic channel can be characterised by two frequency-dependent parameters: the reverberation time and the direct-to-reverberant energy ratio. This paper presents an algorithm for blindly determining these parameters from a single-channel speech signal. The algorithm uses an…

Cited by 0SourceScholar
2015

Speaker change detection and speaker diarization using spatial information

ICASSP 2015accepted

In this paper, we present a novel speaker change detection and speaker diarization algorithm using spatial information in the form of features derived from estimated Room Impulse Response (RIR)s. A blind system identification approach is used to obtain an estimate of the RIRs, from which the C5 feat…

Cited by 13SourceScholar