← Search

Mark J. F. Gales

26 accepted papers

2024

Towards End-to-End Spoken Grammatical Error Correction

ICASSP 2024accepted

Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, wi…

Cited by 0SourceScholar
2023

Ensemble Prosody Prediction For Expressive Speech Synthesis

ICASSP 2023accepted

Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Yet, in evaluations it is generally found that no single model is preferred for al…

Cited by 4SourceScholar
2023

Logit-based ensemble distribution distillation for robust autoregressive sequence uncertainties

UAI 2023poster

Efficiently and reliably estimating uncertainty is an important objective in deep learning. It is especially pertinent to autoregressive sequence tasks, where training and inference costs are typically very high. However, existing research has predominantly focused on tasks with static data such as…

Cited by 5SourcePDFScholar
2021

Analysing Bias in Spoken Language Assessment Using Concept Activation Vectors

ICASSP 2021accepted

A significant concern with deep learning based approaches is that they are difficult to interpret, which means detecting bias in network predictions can be challenging. Concept Activation Vectors (CAVs) have been proposed to address this problem. These use representations - perturbations of activati…

Cited by 0SourceScholar
2021

Ensemble Distillation Approaches for Grammatical Error Correction

ICASSP 2021accepted

Ensemble approaches are commonly used techniques to improving a system by combining multiple model predictions. Additionally these schemes allow the uncertainty, as well as the source of the uncertainty, to be derived for the prediction. Unfortunately these benefits come at a computational and memor…

Cited by 0SourceScholar
2020

Confidence Estimation for Black Box Automatic Speech Recognition Systems Using Lattice Recurrent Neural Networks

ICASSP 2020accepted

Recently, there has been growth in providers of speech transcription services enabling others to leverage technology they would not normally be able to use. As a result, speech-enabled solutions have become commonplace. Their success critically relies on the quality, accuracy, and reliability of the…

Cited by 0SourceScholar
2019

Automatic Grammatical Error Detection of Non-native Spoken Learner English

ICASSP 2019accepted

Automatic language assessment and learning systems are required to support the global growth in English language learning. They need to be able to provide reliable and meaningful feedback to help learners develop their skills. This paper considers the question of detecting "grammatical" errors in no…

Cited by 0SourceScholar
2019

Bi-directional Lattice Recurrent Neural Networks for Confidence Estimation

ICASSP 2019accepted

The standard approach to mitigate errors made by an automatic speech recognition system is to use confidence scores associated with each predicted word. In the simplest case, these scores are word posterior probabilities whilst more complex schemes utilise bi-directional recurrent neural network (Bi…

Cited by 0SourceScholar
2018

Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription

ICASSP 2018accepted

State-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to the spoken form is complicated. However, in recent years the representational powe…

Cited by 0SourceScholar
2017

Morph-to-word transduction for accurate and efficient automatic speech recognition and keyword search

ICASSP 2017accepted

Word units are a popular choice in statistical language modelling. For inflective and agglutinative languages this choice may result in a high out of vocabulary rate. Subword units, such as morphs, provide an interesting alternative to words. These units can be derived in an unsupervised fashion and…

Cited by 0SourceScholar
2017

Recurrent neural network language models for keyword search

ICASSP 2017accepted

Recurrent neural network language models (RNNLMs) have becoming increasingly popular in many applications such as automatic speech recognition (ASR). Significant performance improvements in both perplexity and word error rate over standard n-gram LMs have been widely reported on ASR tasks. In contra…

Cited by 0SourceScholar
2017

Stimulated training for automatic speech recognition and keyword search in limited resource conditions

ICASSP 2017accepted

Training neural network acoustic models on limited quantities of data is a challenging task. A number of techniques have been proposed to improve generalisation. This paper investigates one such technique called stimulated training. It enables standard criteria such as cross-entropy to enforce spati…

Cited by 10SourceScholar
2016

CUED-RNNLM - An open-source toolkit for efficient training and evaluation of recurrent neural network language models

ICASSP 2016accepted

In recent years, recurrent neural network language models (RNNLMs) have become increasingly popular for a range of applications including speech recognition. However, the training of RNNLMs is computationally expensive, which limits the quantity of data, and size of network, that can be used. In ord…

Cited by 0SourceScholar
2016

Combining i-vector representation and structured neural networks for rapid adaptation

ICASSP 2016accepted

Rapid adaptation of deep neural networks (DNNs) with limited unsupervised data remains a significant challenge. This paper investigates the combination of two schemes that have been proposed to address this problem: i-vector representations and multi-basis adaptive neural networks (MBANNs). Two appr…

Cited by 0SourceScholar
2016

Improved DNN-based segmentation for multi-genre broadcast audio

ICASSP 2016accepted

Automatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for…

Cited by 0SourceScholar
2015

Improving multiple-crowd-sourced transcriptions using a speech recogniser

ICASSP 2015accepted

This paper introduces a method to produce high-quality transcriptions of speech data from only two crowd-sourced transcriptions. These transcriptions, produced cheaply by people on the Internet, for example through Amazon Mechanical Turk, are often of low quality. Often, multiple crowd-sourced trans…

Cited by 23SourceScholar
2015

Improving the training and evaluation efficiency of recurrent neural network language models

ICASSP 2015accepted

Recurrent neural network language models (RNNLMs) are becoming increasingly popular for speech recognition. Previously, we have shown that RNNLMs with a full (non-classed) output layer (F-RNNLMs) can be trained efficiently using a GPU giving a large reduction in training time over conventional class…

Cited by 0SourceScholar
2015

Recurrent neural network language model training with noise contrastive estimation for speech recognition

ICASSP 2015accepted

In recent years recurrent neural network language models (RNNLMs) have been successfully applied to a range of tasks including speech recognition. However, an important issue that limits the quantity of data used, and their possible application areas, is the computational cost in training. A signi??…

Cited by 0SourceScholar
2015

Robust excitation-based features for Automatic Speech Recognition

ICASSP 2015accepted

In this paper we investigate the use of noise-robust features characterizing the speech excitation signal as complementary features to the usually considered vocal tract based features for Automatic Speech Recognition (ASR). The proposed Excitation-based Features (EBF) are tested in a state-of-the-a…

Cited by 0SourceScholar