← Search

Peter Bell

24 accepted papers

2026

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

ICLR 2026oral

Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but rarely validated against subjective ones. Both kinds of metrics are challenged by…

Cited by 0SourcecodeScholar
2025

LLM-Personalize: Aligning LLM Planners with Human Preferences via Reinforced Self-Training for Housekeeping Robots

COLING 2025main

Large language models (LLMs) have shown significant potential for robotics applications, particularly task planning, by harnessing their language comprehension and text generation capabilities. However, in applications such as household robotics, a critical gap remains in the personalization of thes…

Cited by 18SourcePDFScholar
2025

Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction

ICASSP 2025accepted

Annotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study on this topic, beginning with the proposal of novel prompts…

Cited by 0SourceScholar
2025

Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling

ICASSP 2025accepted

The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling met…

Cited by 0SourceScholar
2025

Spoken Document Retrieval for an Unwritten Language: A Case Study on Gormati

EMNLP 2025

Speakers of unwritten languages have the potential to benefit from speech-based automatic information retrieval systems. This paper proposes a speech embedding technique that facilitates such a system that we can be used in a zero-shot manner on the target language. After conducting development expe

Cited by 0SourcePDFScholar
2024

Bootstrap Predictive Coding: Investigating a Non-Contrastive Self-Supervised Learning Approach

ICASSP 2024accepted

Self-supervised learning methods (SSL) have seen wide popularity for speech representation learning. Early methods, such as wav2vec, were causal, whilst more recent approaches, notably wav2vec 2.0 and data2vec, have employed masking strategies together with a Transformer architecture. Many SSL metho…

Cited by 0SourceScholar
2024

Can We Trust Explainable AI Methods on ASR? An Evaluation on Phoneme Recognition

ICASSP 2024accepted

Explainable AI (XAI) techniques have been widely used to help explain and understand the output of deep learning models in fields such as image classification and Natural Language Processing. Interest in using XAI techniques to explain deep learning-based Automatic Speech Recognition (ASR) is emergi…

Cited by 0SourceScholar
2023

Efficient Intelligibility Evaluation Using Keyword Spotting: A Study on Audio-Visual Speech Enhancement

ICASSP 2023accepted

We propose a new method for human speech intelligibility evaluation based on keyword spotting. In this method, participants play a stimulus and select the word they hear from a close set of alternatives. To find which sentence to use, the target word, and alternatives we mine a large set of stimuli…

Cited by 0SourceScholar
2023

Multimodal Dyadic Impression Recognition via Listener Adaptive Cross-Domain Fusion

ICASSP 2023accepted

As a sub-branch of affective computing, impression recognition, e.g., perception of speaker characteristics such as warmth or competence, is potentially a critical part of both human-human conversations and spoken dialogue systems. Most research has studied impressions only from the behaviors expres…

Cited by 0SourceScholar
2023

The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR

ICASSP 2023accepted

English is the most widely spoken language in the world, used daily by millions of people as a first or second language in many different contexts. As a result, there are many varieties of English. Although the great many advances in English automatic speech recognition (ASR) over the past decades,…

Cited by 0SourceScholar
2021

Train Your Classifier First: Cascade Neural Networks Training from Upper Layers to Lower Layers

ICASSP 2021accepted

Although the lower layers of a deep neural network learn features which are transferable across datasets, these layers are not transferable within the same dataset. That is, in general, freezing the trained feature extractor (the lower layers) and retraining the classifier (the upper layers) on the…

Cited by 0SourceScholar
2020

Cross Lingual Transfer Learning for Zero-Resource Domain Adaptation

ICASSP 2020accepted

We propose a method for zero-resource domain adaptation of DNN acoustic models, for use in low-resource situations where the only in-language training data available may be poorly matched to the intended target domain. Our method uses a multi-lingual model in which several DNN layers are shared betw…

Cited by 0SourceScholar
2019

On the Usefulness of Statistical Normalisation of Bottleneck Features for Speech Recognition

ICASSP 2019accepted

DNNs play a major role in the state-of-the-art ASR systems. They can be used for extracting features and building probabilistic models for acoustic and language modelling. Despite their huge practical success, the level of theoretical understanding has remained shallow. This paper investigates DNNs…

Cited by 2SourceScholar
2017

Sequence-to-sequence models for punctuated transcription combining lexical and acoustic features

ICASSP 2017accepted

In this paper we present an extension of our previously described neural machine translation based system for punctuated transcription. This extension allows the system to map from per frame acoustic features to word level representations by replacing the traditional encoder in the encoder-decoder a…

Cited by 0SourceScholar
2015

Regularization of context-dependent deep neural networks with context-independent multi-task training

ICASSP 2015accepted

The use of context-dependent targets has become standard in hybrid DNN systems for automatic speech recognition. However, we argue that despite the use of state-tying, optimising to context-dependent targets can lead to over-fitting, and that discriminating between arbitrary tied context-dependent t…

Cited by 0SourceScholar