← Search

Patrick A. Naylor

32 accepted papers

2024

Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks

ICASSP 2024accepted

Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an enco…

Cited by 0SourceScholar
2024

Speech Enhancement in Hearing Aids Using Target Speech Presence Estimation Based on a Delayed Remote Microphone Signal

ICASSP 2024accepted

Speech enhancement in hearing aids (HAs) can take advantage of a wireless remote microphone (RM) having a better signal-to-noise ratio than the HA microphones. However, using the RM effectively is complicated by the time delay between the acoustic and wireless signals. Methods in the literature assu…

Cited by 0SourceScholar
2023

Graph Neural Networks for Sound Source Localization on Distributed Microphone Networks

ICASSP 2023accepted

Distributed Microphone Arrays (DMAs) present many challenges with respect to centralized microphone arrays. An important requirement of applications on these arrays is handling a variable number of input channels. We consider the use of Graph Neural Networks (GNNs) as a solution to this challenge. W…

Cited by 0SourceScholar
2023

Subspace Hybrid Beamforming for Head-Worn Microphone Arrays

ICASSP 2023accepted

A two-stage multi-channel speech enhancement method is proposed which consists of a novel adaptive beamformer, Hybrid Minimum Variance Distortionless Response (MVDR), Isotropic-MVDR (Iso), and a novel multi-channel spectral Principal Components Analysis (PCA) denoising. In the first stage, the Hybri…

Cited by 0SourceScholar
2023

The MBSTOI Binaural Intelligibility Metric Using a Close-Talking Microphone Reference

ICASSP 2023accepted

Intelligibility metrics are a fast way to determine how comprehensible a target signal is in a noisy situation. Most metrics however rely on having a clean reference signal for computation and are not adapted to live recordings. In this paper the deep correlation modified binaural short time objecti…

Cited by 0SourceScholar
2022

Spatial Processing Front-End for Distant ASR Exploiting Self-Attention Channel Combinator

ICASSP 2022accepted

We present a novel multi-channel front-end based on channel shortening with the Weighted Prediction Error (WPE) method followed by a fixed MVDR beamformer used in combination with a recently proposed self-attention-based channel combination (SACC) scheme, for tackling the distant ASR problem. We sho…

Cited by 7SourceScholar
2021

Multichannel Overlapping Speaker Segmentation Using Multiple Hypothesis Tracking Of Acoustic And Spatial Features

ICASSP 2021accepted

An essential part of any diarization system is the task of speaker segmentation which is important for many applications including speaker indexing and automatic speech recognition (ASR) in multi-speaker environments. Segmentation of overlapping speech has recently been a key focus of this work. In…

Cited by 0SourceScholar
2021

Polynomial Matrix Eigenvalue Decomposition of Spherical Harmonics for Speech Enhancement

ICASSP 2021accepted

Speech enhancement algorithms using polynomial matrix eigenvalue decomposition (PEVD) have been shown to be effective for noisy and reverberant speech. However, these algorithms do not scale well in complexity with the number of channels used in the processing. For a spherical microphone array sampl…

Cited by 0SourceScholar
2021

Processing Pipelines for Efficient, Physically-Accurate Simulation of Microphone Array Signals in Dynamic Sound Scenes

ICASSP 2021accepted

Multichannel acoustic signal processing is predicated on the fact that the interchannel relationships between the received signals can be exploited to infer information about the acoustic scene. Recently there has been increasing interest in algorithms which are applicable in dynamic scenes, where t…

Cited by 2SourceScholar
2019

Second Order Sequential Best Rotation Algorithm with Householder Reduction for Polynomial Matrix Eigenvalue Decomposition

ICASSP 2019accepted

The Second-order Sequential Best Rotation (SBR2) algorithm, used for Eigenvalue Decomposition (EVD) on para-Hermitian polynomial matrices typically encountered in wideband signal processing applications like multichannel Wiener filtering and channel coding, involves a series of delay and rotation op…

Cited by 0SourceScholar
2019

Speaker Change Detection Using Fundamental Frequency with Application to Multi-talker Segmentation

ICASSP 2019accepted

This paper shows that time varying pitch properties can be used advantageously within the segmentation step of a multi-talker diarization system. First a study is conducted to verify that changes in pitch are strong indicators of changes in the speaker. It is then highlighted that an individual's pi…

Cited by 0SourceScholar
2018

Acoustic Analysis and Assessment of the Knee in Osteoarthritis During Walking

ICASSP 2018accepted

We examine the relation between the sounds emitted by the knee joint during walking and its condition, with particular focus on osteoarthritis, and investigate their potential for noninvasive detection of knee pathology. We present a comparative analysis of several features and evaluate their discri…

Cited by 0SourceScholar
2018

Joint Source Localization and Dereverberation by Sound Field Interpolation Using Sparse Regularization

ICASSP 2018accepted

In this paper, source localization and dereverberation are formulated jointly as an inverse problem. The inverse problem consists in the interpolation of the sound field measured by a set of microphones by matching the recorded sound pressure with that of a particular acoustic model. This model is b…

Cited by 0SourceScholar
2018

Room Identification Using Frequency Dependence of Spectral Decay Statistics

ICASSP 2018accepted

A method for room identification is proposed based on the reverberation properties of multichannel speech recordings. The approach exploits the dependence of spectral decay statistics on the reverberation time of a room. The average negative-side variance within 1/3-octave bands is proposed as the i…

Cited by 4SourceScholar
2017

A dynamic programming approach for automatic stride detection and segmentation in acoustic emission from the knee

ICASSP 2017accepted

We study the acquisition and analysis of sounds generated by the knee during walking with particular focus on the effects due to osteoarthritis. Reliable contact instant estimation is essential for stride synchronous analysis. We present a dynamic programming based algorithm for automatic estimation…

Cited by 0SourceScholar
2017

Discriminative feature domains for reverberant acoustic environments

ICASSP 2017accepted

Several speech processing and audio data-mining applications rely on a description of the acoustic environment as a feature vector for classification. The discriminative properties of the feature domain play a crucial role in the effectiveness of these methods. In this work, we consider three enviro…

Cited by 0SourceScholar
2017

Frequency-domain under-modelled blind system identification based on cross power spectrum and sparsity regularization

ICASSP 2017accepted

In room acoustics, under-modelled multichannel blind system identification (BSI) aims to estimate the early part of the room impulse responses (RIRs), and it can be widely used in applications such as speaker localization, room geometry identification and beamforming based speech dereverberation. In…

Cited by 0SourceScholar
2017

Improving the perceptual quality of ideal binary masked speech

ICASSP 2017accepted

It is known that applying a time-frequency binary mask to very noisy speech can improve its intelligibility but results in poor perceptual quality. In this paper we propose a new approach to applying a binary mask that combines the intelligibility gains of conventional binary masking with the percep…

Cited by 0SourceScholar
2017

Measuring, modelling and predicting perceived reverberation

ICASSP 2017accepted

This paper investigates the relationship between the perceived level of reverberation and parameters measured from the room impulse response (RIR), as well as the design of an instrumental measure that predicts this perceived level. We first present the results of an experimental listening test cond…

Cited by 0SourceScholar
2017

Multiple source localization using Estimation Consistency in the Time-Frequency domain

ICASSP 2017accepted

The extraction of multiple Direction-of-Arrival (DoA) information from estimated spatial spectra can be challenging when such spectra are noisy or the sources are adjacent. Smoothing or clustering techniques are typically used to remove the effect of noise or irregular peaks in the spatial spectra.…

Cited by 0SourceScholar
2017

Robust spherical harmonic domain interpolation of spatially sampled array manifolds

ICASSP 2017accepted

Accurate interpolation of the array manifold is an important first step for the acoustic simulation of rapidly moving microphone arrays. Spherical harmonic domain interpolation has been proposed and well studied in the context of head-related transfer functions but has focussed on perceptual, rather…

Cited by 0SourceScholar
2017

Source tracking using moving microphone arrays for robot audition

ICASSP 2017accepted

Intuitive spoken dialogues are a prerequisite for human-robot interaction. In many practical situations, robots must be able to identify and focus on sources of interest in the presence of interfering speakers. Techniques such as spatial filtering and blind source separation are therefore often used…

Cited by 0SourceScholar
2016

3D acoustic source localization in the spherical harmonic domain based on optimized grid search

ICASSP 2016accepted

An approach for 3D source localization using a spherical microphone array is proposed that gives improved accuracy compared to intensity-based methods. First order spherical harmonics are first used to obtain an initial approximate localization result and then the initial result is improved based on…

Cited by 15SourceScholar
2016

Acoustic simultaneous localization and mapping (A-SLAM) of a moving microphone array and its surrounding speakers

ICASSP 2016accepted

Acoustic scene mapping creates a representation of positions of audio sources such as talkers within the surrounding environment of a microphone array. By allowing the array to move, the acoustic scene can be explored in order to improve the map. Furthermore, the spatial diversity of the kinematic a…

Cited by 46SourceScholar
2016

Perceptual and instrumental evaluation of the perceived level of reverberation

ICASSP 2016accepted

Perceptual measures are usually considered more reliable than instrumental measures for evaluating the perceived level of reverberation. However, such measures are costly in both time and money, and, due to variations in stimuli or assessors, the resulting data is not always statistically significan…

Cited by 13SourceScholar
2015

Direct-to-Reverberant Ratio estimation using a null-steered beamformer

ICASSP 2015accepted

Reverberation affects the quality and intelligibility of distant speech recorded in a room. Direct-to-Reverberant Ratio (DRR) is a useful measure for assessing the acoustic configuration and can be used to inform dereverberation algorithms. We describe a novel DRR estimation algorithm applicable whe…

Cited by 0SourceScholar
2015

Single-channel blind estimation of reverberation parameters

ICASSP 2015accepted

The reverberation of an acoustic channel can be characterised by two frequency-dependent parameters: the reverberation time and the direct-to-reverberant energy ratio. This paper presents an algorithm for blindly determining these parameters from a single-channel speech signal. The algorithm uses an…

Cited by 0SourceScholar
2015

Speaker change detection and speaker diarization using spatial information

ICASSP 2015accepted

In this paper, we present a novel speaker change detection and speaker diarization algorithm using spatial information in the form of features derived from estimated Room Impulse Response (RIR)s. A blind system identification approach is used to obtain an estimate of the RIRs, from which the C5 feat…

Cited by 13SourceScholar