← Search

Philip C. Woodland

27 accepted papers

2025

Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition

ICASSP 2025accepted

While large language models (LLMs) have been applied to automatic speech recognition (ASR), the task of making the model streamable remains a challenge. This paper proposes a novel model architecture, Transducer-Llama, that integrates LLMs into a Factorized Transducer (FT) model, naturally enabling…

Cited by 0SourceScholar
2024

Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation

ICASSP 2024accepted

Foundation models have shown superior performance for speech emotion recognition (SER). However, given the limited data in emotion corpora, finetuning all parameters of large pre-trained models for SER can be both resource-intensive and susceptible to overfitting. This paper investigates parameter-e…

Cited by 0SourceScholar
2023

End-to-End Spoken Language Understanding with Tree-Constrained Pointer Generator

ICASSP 2023accepted

End-to-end spoken language understanding (SLU) suffers from the long-tail word problem. This paper exploits contextual biasing, a technique to improve the speech recognition of rare words, in end-to-end SLU systems. Specifically, a tree-constrained pointer generator (TCPGen), a powerful and efficien…

Cited by 0SourceScholar
2023

Spectral Clustering-Aware Learning of Embeddings for Speaker Diarisation

ICASSP 2023accepted

In speaker diarisation, speaker embedding extraction models often suffer from the mismatch between their training loss functions and the speaker clustering method. In this paper, we propose the method of spectral clustering-aware learning of embeddings (SCALE) to address the mismatch. Specifically,…

Cited by 1SourceScholar
2022

Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition

ICASSP 2022accepted

As end-to-end automatic speech recognition (ASR) models reach promising performance, various downstream tasks rely on good confidence estimators for these systems. Recent research has shown that model-based confidence estimators have a significant advantage over using the output softmax probabilitie…

Cited by 16SourceScholar
2022

Knowledge Distillation for Neural Transducers from Large Self-Supervised Pre-Trained Models

ICASSP 2022accepted

Self-supervised pre-training is an effective approach to leveraging a large amount of unlabelled data to reduce word error rates (WERs) of automatic speech recognition (ASR) systems. Since it is impractical to use large pre-trained models for many real-world ASR applications, it is desirable to have…

Cited by 29SourceScholar
2021

Confidence Estimation for Attention-Based Sequence-to-Sequence Models for Speech Recognition

ICASSP 2021accepted

For various speech-related tasks, confidence scores from a speech recogniser are a useful measure to assess the quality of transcriptions. In traditional hidden Markov model-based automatic speech recognition (ASR) systems, confidence scores can be reliably obtained from word posteriors in decoding…

Cited by 0SourceScholar
2021

Emotion Recognition by Fusing Time Synchronous and Time Asynchronous Representations

ICASSP 2021accepted

In this paper, a novel two-branch neural network model structure is proposed for multimodal emotion recognition, which consists of a time synchronous branch (TSB) and a time asynchronous branch (TAB). To capture correlations between each word and its acoustic realisation, the TSB combines speech and…

Cited by 0SourceScholar
2021

Transformer Language Models with LSTM-Based Cross-Utterance Information Representation

ICASSP 2021accepted

The effective incorporation of cross-utterance information has the potential to improve language models (LMs) for automatic speech recognition (ASR). To extract more powerful and robust cross-utterance representations for the Transformer LM (TLM), this paper proposes the R-TLM which uses hidden stat…

Cited by 0SourceScholar
2018

Improved Tdnns Using Deep Kernels and Frequency Dependent Grid-RNNS

ICASSP 2018accepted

Time delay neural networks (TDNNs) are an effective acoustic model for large vocabulary speech recognition. The strength of the model can be attributed to its ability to effectively model long temporal contexts. However, current TDNN models are relatively shallow, which limits the modelling capabili…

Cited by 23SourceScholar
2017

Joint optimisation of tandem systems using Gaussian mixture density neural network discriminative sequence training

ICASSP 2017accepted

The use of deep neural networks (DNNs) for feature extraction and Gaussian mixture models (GMMs) for acoustic modelling is often termed a tandem system configuration and can be viewed as a Gaussian mixture density neural network (MDNN). Compared to the direct use of DNN output probabilities in the a…

Cited by 0SourceScholar
2016

CUED-RNNLM - An open-source toolkit for efficient training and evaluation of recurrent neural network language models

ICASSP 2016accepted

In recent years, recurrent neural network language models (RNNLMs) have become increasingly popular for a range of applications including speech recognition. However, the training of RNNLMs is computationally expensive, which limits the quantity of data, and size of network, that can be used. In ord…

Cited by 0SourceScholar
2016

DNN speaker adaptation using parameterised sigmoid and ReLU hidden activation functions

ICASSP 2016accepted

This paper investigates the use of parameterised sigmoid and rectified linear unit (ReLU) hidden activation functions in deep neural network (DNN) speaker adaptation. The sigmoid and ReLU parameterisation schemes from a previous study for speaker independent (SI) training are used. An adaptive linea…

Cited by 0SourceScholar
2016

Improved DNN-based segmentation for multi-genre broadcast audio

ICASSP 2016accepted

Automatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for…

Cited by 0SourceScholar
2015

Improving the training and evaluation efficiency of recurrent neural network language models

ICASSP 2015accepted

Recurrent neural network language models (RNNLMs) are becoming increasingly popular for speech recognition. Previously, we have shown that RNNLMs with a full (non-classed) output layer (F-RNNLMs) can be trained efficiently using a GPU giving a large reduction in training time over conventional class…

Cited by 0SourceScholar
2015

Recurrent neural network language model training with noise contrastive estimation for speech recognition

ICASSP 2015accepted

In recent years recurrent neural network language models (RNNLMs) have been successfully applied to a range of tasks including speech recognition. However, an important issue that limits the quantity of data used, and their possible application areas, is the computational cost in training. A signi??…

Cited by 0SourceScholar