← Search

Jon Barker

17 accepted papers

2024

Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory Models

ICASSP 2024accepted

Neural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work…

Cited by 0SourceScholar
2024

The 2nd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing Aid Intelligibility Prediction

ICASSP 2024accepted

This paper reports on the design and outcomes of the 2nd Clarity Prediction Challenge (CPC2) for predicting the intelligibility of hearing aid processed signals heard by individuals with a hearing impairment. The challenge was designed to promote new approaches for estimating the intelligibility of…

Cited by 0SourceScholar
2023

Overview of the 2023 ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids

ICASSP 2023accepted

This paper reports on the design and outcomes of the ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids. The scenario was a listener attending to a target speaker in a noisy, domestic environment. There were multiple interferers and head rotation by the listener. The challenge extended…

Cited by 0SourceScholar
2023

The 2nd Clarity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and Outcomes

ICASSP 2023accepted

This paper reports on the design and outcomes of the 2nd Clarity Enhancement Challenge (CEC2), a challenge for stimulating novel approaches to hearing-aid speech intelligibility enhancement. The challenge was for a listener attending to a target speaker in a noisy, domestic environment. The challeng…

Cited by 0SourceScholar
2022

Auditory-Based Data Augmentation for end-to-end Automatic Speech Recognition

ICASSP 2022accepted

End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired front-ends have also demonstrated improvement for automatic spee…

Cited by 0SourceScholar
2022

Improved Simulation of Realistically-Spatialised Simultaneous Speech Using Multi-Camera Analysis in The Chime-5 Dataset

ICASSP 2022accepted

Room simulation is an essential tool in the development of distant microphone ASR and source separation. However, most commonly used simulated datasets adopt uninformed and potentially unrealistic speaker location distributions. In earlier work, we analysed a 50-hour audio-visual dataset of multipar…

Cited by 0SourceScholar
2022

Multi-Modal Acoustic-Articulatory Feature Fusion For Dysarthric Speech Recognition

ICASSP 2022accepted

Building automatic speech recognition (ASR) systems for speakers with dysarthria is a very challenging task. Although multi-modal ASR has received increasing attention recently, incorporating real articulatory data with acoustic features has not been widely explored in the dysarthric speech communit…

Cited by 0SourceScholar
2021

Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism

ICASSP 2021accepted

In this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on an improved multi-channel time-domain speech separation network which employs spe…

Cited by 0SourceScholar
2020

Exploring Appropriate Acoustic and Language Modelling Choices for Continuous Dysarthric Speech Recognition

ICASSP 2020accepted

There has been much recent interest in building continuous speech recognition systems for people with severe speech impairments, e.g., dysarthria. However, the datasets that are commonly used are typically designed for tasks other than ASR development, or they contain only isolated words. As such, t…

Cited by 0SourceScholar
2020

On End-to-end Multi-channel Time Domain Speech Separation in Reverberant Environments

ICASSP 2020accepted

This paper introduces a new method for multi-channel time domain speech separation in reverberant environments. A fully-convolutional neural network structure has been used to directly separate speech from multiple microphone recordings, with no need of conventional spatial feature extraction. To re…

Cited by 57SourceScholar
2020

Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech Recognition

ICASSP 2020accepted

This paper presents an improved transfer learning framework applied to robust personalised speech recognition models for speakers with dysarthria. As the baseline of transfer learning, a state-of-the-art CNN-TDNN-F ASR acoustic model trained solely on source domain data is adapted onto the target do…

Cited by 0SourceScholar
2019

Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech Recognition

ICASSP 2019accepted

Improving the accuracy of personalised speech recognition for speakers with dysarthria is a challenging research field. In this paper, we explore an approach that non-linearly modifies speech tempo to reduce mismatch between typical and atypical speech. Speech tempo analysis at the phonetic level is…

Cited by 0SourceScholar
2018

Exploring the Use of Group Delay for Generalised VTS Based Noise Compensation

ICASSP 2018accepted

In earlier work we studied the effect of statistical normalisation for phase-based features and observed it leads to a significant robustness improvement. This paper explores the extension of the generalised Vector Taylor Series (gVTS) noise compensation approach to the group delay (GD) domain. We d…

Cited by 0SourceScholar
2018

SDC-Net: Video prediction using spatially-displaced convolution

ECCV 2018poster

We present an approach for high-resolution video frame prediction by conditioning on both past frames and past optical flows. Previous approaches rely on resampling past frames, guided by a learned future optical flow, or on direct generation of pixels. Resampling based on flow is insufficient becau…

2017

Statistical normalisation of phase-based feature representation for robust speech recognition

ICASSP 2017accepted

In earlier work we have proposed a source-filter decomposition of speech through phase-based processing. The decomposition leads to novel speech features that are extracted from the filter component of the phase spectrum. This paper analyses this spectrum and the proposed representation by evaluatin…

Cited by 0SourceScholar