← Search

Herman Kamper

15 accepted papers

2025

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model

ICASSP 2025accepted

Codec-based text-to-speech (TTS) models have shown impressive quality with zero-shot voice cloning abilities. However, they often struggle with more expressive references or complex text inputs. We present MARS6, a robust encoder-decoder transformer for rapid, expressive TTS. MARS6 is built on recen…

Cited by 0SourceScholar
2025

Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives

ICASSP 2025accepted

We develop automatic speech recognition (ASR) systems for stories told by Afrikaans and isiXhosa preschool children. Oral narratives provide a way to assess children’s language development before they learn to read. We consider a range of prior child-speech ASR strategies to determine which is best…

Cited by 0SourceScholar
2025

Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming

ICASSP 2025accepted

We look at the long-standing problem of segmenting unlabeled speech into word-like segments and clustering these into a lexicon. Several previous methods use a scoring model coupled with dynamic programming to find an optimal segmentation. Here we propose a much simpler strategy: we predict word bou…

Cited by 0SourceScholar
2022

A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion

ICASSP 2022accepted

The goal of voice conversion is to transform source speech into a target voice, keeping the content unchanged. In this paper, we focus on self-supervised representation learning for voice conversion. Specifically, we compare discrete and soft speech units as input features. We find that discrete rep…

Cited by 0SourceScholar
2020

Cross-Lingual Topic Prediction For Speech Using Translations

ICASSP 2020accepted

Given a large amount of unannotated speech in a low-resource language, can we classify the speech utterances by topicƒ We consider this question in the setting where a small amount of speech in the low-resource language is paired with text translations in a high-resource language. We develop an effe…

Cited by 0SourceScholar
2020

Multilingual Acoustic Word Embedding Models for Processing Zero-resource Languages

ICASSP 2020accepted

Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. In settings where unlabelled speech is the only available resource, such embeddings can be used in "zero-resource" speech search, indexing and discovery systems. Here we propose to train a single super…

Cited by 0SourceScholar
2019

Semantic Query-by-example Speech Search Using Visual Grounding

ICASSP 2019accepted

A number of recent studies have started to investigate how speech systems can be trained on untranscribed speech by leveraging accompanying images at training time. Examples of tasks include keyword prediction and within- and across-mode retrieval. Here we consider how such models can be used for qu…

Cited by 0SourceScholar
2018

Critical initialisation for deep signal propagation in noisy rectifier neural networks

NeurIPS 2018poster

Stochastic regularisation is an important weapon in the arsenal of a deep learning practitioner. However, despite recent theoretical advances, our understanding of how noise influences signal propagation in deep neural networks remains limited. By extending recent work based on mean field theory, we…

2018

Phoneme Based Embedded Segmental K-Means for Unsupervised Term Discovery

ICASSP 2018accepted

Identifying and grouping the frequently occurring word-like patterns from raw acoustic waveforms is an important task in the zero resource speech processing. Embedded segmental K-means (ES-KMeans) discovers both the word boundaries and the word types from raw data. Starting from an initial set of su…

Cited by 0SourceScholar
2017

Weakly supervised spoken term discovery using cross-lingual side information

ICASSP 2017accepted

Recent work on unsupervised term discovery (UTD) aims to identify and cluster repeated word-like units from audio alone. These systems are promising for some very low-resource languages where transcribed audio is unavailable, or where no written form of the language exists. However, in some cases it…

Cited by 0SourceScholar
2016

Deep convolutional acoustic word embeddings using word-pair side information

ICASSP 2016accepted

Recent studies have been revisiting whole words as the basic modelling unit in speech recognition and query applications, instead of phonetic units. Such whole-word segmental systems rely on a function that maps a variable-length speech segment to a vector in a fixed-dimensional space; the resulting…

Cited by 0SourceScholar
2015

Unsupervised neural network based feature extraction using weak top-down constraints

ICASSP 2015accepted

Deep neural networks (DNNs) have become a standard component in supervised ASR, used in both data-driven feature extraction and acoustic modelling. Supervision is typically obtained from a forced alignment that provides phone class targets, requiring transcriptions and pronunciations. We propose a n…

Cited by 0SourceScholar