← Search

Thomas Hain

26 accepted papers

2025

Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text

ICASSP 2025accepted

Word error rate (WER) estimation aims to evaluate the quality of an automatic speech recognition (ASR) system’s output without requiring ground-truth labels. This task has gained increasing attention as advanced ASR systems are trained on large amounts of data. In this context, the computational eff…

Cited by 0SourceScholar
2024

Automatic Speech Recognition System-Independent Word Error Rate Estimation

COLING 2024main

Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition (ASR) systems. In many applications, it is of interest to estimate WER given a pair of a speech utterance and a transcript. Previous work on WER estimation focused on building mo…

2024

Combining Conformer and Dual-Path-Transformer Networks for Single Channel Noisy Reverberant Speech Separation

ICASSP 2024accepted

Separation of overlapping speakers remains an active area of speech technology research. Many deep neural network (DNN) separation models propose modelling local and global temporal context separately using alternating DNN layers. Two such models are SepFormer and TD-Conformer. The largest configura…

Cited by 0SourceScholar
2024

Multi-CMGAN+/+: Leveraging Multi-Objective Speech Quality Metric Prediction for Speech Enhancement

ICASSP 2024accepted

Neural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on artificially created labelled training data such that the ne…

Cited by 0SourceScholar
2024

Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory Models

ICASSP 2024accepted

Neural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work…

Cited by 0SourceScholar
2024

Progressive Unsupervised Domain Adaptation for ASR Using Ensemble Models and Multi-Stage Training

ICASSP 2024accepted

In Automatic Speech Recognition (ASR), teacher-student (T/S) training has shown to perform well for domain adaptation with small amount of training data. However, adaption without ground-truth labels is still challenging. A previous study has shown the effectiveness of using ensemble teacher models…

Cited by 0SourceScholar
2023

Deformable Temporal Convolutional Networks for Monaural Noisy Reverberant Speech Separation

ICASSP 2023accepted

Speech separation models are used for isolating individual speakers in many speech processing applications. Deep learning models have been shown to lead to state-of-the-art (SOTA) results on a number of speech separation benchmarks. One such class of models known as temporal convolutional networks (…

Cited by 0SourceScholar
2023

Perceive and Predict: Self-Supervised Speech Representation Based Loss Functions for Speech Enhancement

ICASSP 2023accepted

Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the deepest or final outputs of self supervised speech representation models, rather…

Cited by 0SourceScholar
2023

Towards Domain Generalisation in ASR with Elitist Sampling and Ensemble Knowledge Distillation

ICASSP 2023accepted

Knowledge distillation (KD) has widely been used for model compression and domain adaptation for speech applications. In the presence of multiple teachers, knowledge can easily be transferred to the student by averaging the models output. However, previous research shows that the student do not adap…

Cited by 3SourceScholar
2022

Unsupervised Data Selection for Speech Recognition with Contrastive Loss Ratios

ICASSP 2022accepted

This paper proposes an unsupervised data selection method by using a submodular function based on contrastive loss ratios of target and training data sets. A model using a contrastive loss function is trained on both sets. Then the ratio of frame-level losses for each model is used by a submodular f…

Cited by 0SourceScholar
2021

Multiple-Hypothesis CTC-Based Semi-Supervised Adaptation of End-to-End Speech Recognition

ICASSP 2021accepted

This paper proposes an adaptation method for end-to-end speech recognition. In this method, multiple automatic speech recognition (ASR) 1-best hypotheses are integrated in the computation of the connectionist temporal classification (CTC) loss function. The integration of multiple ASR hypotheses hel…

Cited by 0SourceScholar
2021

Towards Low-Resource Stargan Voice Conversion Using Weight Adaptive Instance Normalization

ICASSP 2021accepted

Many-to-many voice conversion with non-parallel training data has seen significant progress in recent years. It is challenging because of lacking of ground truth parallel data. StarGAN-based models have gained attentions because of their efficiency and effective. However, most of the StarGAN-based w…

Cited by 0SourceScholar
2020

H-Vectors: Utterance-Level Speaker Embedding Using a Hierarchical Attention Model

ICASSP 2020accepted

In this paper, a hierarchical attention network is proposed to generate utterance-level embeddings (H-vectors) for speaker identification and verification. Since different parts of an utterance may have different contributions to speaker identities, the use of hierarchical structure aims to learn sp…

Cited by 0SourceScholar
2018

Exploring the Use of Group Delay for Generalised VTS Based Noise Compensation

ICASSP 2018accepted

In earlier work we studied the effect of statistical normalisation for phase-based features and observed it leads to a significant robustness improvement. This paper explores the extension of the generalised Vector Taylor Series (gVTS) noise compensation approach to the group delay (GD) domain. We d…

Cited by 0SourceScholar
2017

Shefce: A Cantonese-English bilingual speech corpus for pronunciation assessment

ICASSP 2017accepted

This paper introduces the development of ShefCE: a Cantonese-English bilingual speech corpus from L2 English speakers in Hong Kong. Bilingual parallel recording materials were chosen from TED online lectures. Script selection were carried out according to bilingual consistency (evaluated using a mac…

Cited by 0SourceScholar
2017

Statistical normalisation of phase-based feature representation for robust speech recognition

ICASSP 2017accepted

In earlier work we have proposed a source-filter decomposition of speech through phase-based processing. The decomposition leads to novel speech features that are extracted from the filter component of the phase spectrum. This paper analyses this spectrum and the proposed representation by evaluatin…

Cited by 0SourceScholar
2016

Groupwise learning for ASR k-best list reranking in spoken language translation

ICASSP 2016accepted

Quality estimation models are used to predict the quality of the output from a spoken language translation (SLT) system. When these scores are used to rerank a k-best list, the rank of the scores is more important than their absolute values. This paper proposes groupwise learning to model this rank.…

Cited by 0SourceScholar
2015

Automatic assessment of English learner pronunciation using discriminative classifiers

ICASSP 2015accepted

This paper presents a novel system for automatic assessment of pronunciation quality of English learner speech, based on deep neural network (DNN) features and phoneme specific discriminative classifiers. DNNs trained on a large corpus of native and non-native learner speech are used to extract phon…

Cited by 0SourceScholar
2015

Quality estimation for asr k-best list rescoring in spoken language translation

ICASSP 2015accepted

Spoken language translation (SLT) combines automatic speech recognition (ASR) and machine translation (MT). During the decoding stage, the best hypothesis produced by the ASR system may not be the best input candidate to the MT system, but making use of multiple sub-optimal ASR results in SLT has be…

Cited by 0SourceScholar