← Search

Erica Cooper

11 accepted papers

2025

Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels

ICASSP 2025accepted

In practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the diction…

Cited by 0SourceScholar
2025

Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores

ICASSP 2025accepted

This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rendering (EPR) model with a fine-tuned neural MIDI synthesiser, our approach directly generates expressive audio performanc…

Cited by 0SourceScholar
2024

Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset

ICASSP 2024accepted

The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech coll…

Cited by 0SourceScholar
2023

Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?

ICASSP 2023accepted

With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS techniques. In this study, we analyze the shortcomings of a TTS-bas…

Cited by 0SourceScholar
2022

Attention Back-End for Automatic Speaker Verification with Multiple Enrollment Utterances

ICASSP 2022accepted

Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment utterances, we propose a novel attention back-end model that can…

Cited by 0SourceScholar
2022

LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech

ICASSP 2022accepted

An effective approach to automatically predict the subjective rating for synthetic speech is to train on a listening test dataset with human-annotated scores. Although each speech sample in the dataset is rated by several listeners, most previous works only used the mean score as the training target…

Cited by 0SourceScholar
2022

On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis

ICASSP 2022accepted

Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a starting point to explore pruning both spectrogram prediction networks and vocoders. We thoroughly investigate the tradeoff…

Cited by 0SourceScholar
2021

How Similar or Different is Rakugo Speech Synthesizer to Professional Performers?

ICASSP 2021accepted

We have been working on speech synthesis for rakugo (a traditional Japanese form of verbal entertainment similar to one-person stand-up comedy) toward speech synthesis that authentically entertains audiences. In this paper, we propose a novel evaluation methodology using synthesized rakugo speech an…

Cited by 0SourceScholar
2021

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

ICASSP 2021accepted

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not generalize well to unseen speakers or content. To alleviate this problem, we have incorporated a speaker encoder and sp…

Cited by 0SourceScholar
2020

Zero-Shot Multi-Speaker Text-To-Speech with State-Of-The-Art Neural Speaker Embeddings

ICASSP 2020accepted

While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zero-shot adaptation to unseen speakers. We investigate multi-speaker modeling for end-to-end text-to-speech synthesis and…

Cited by 0SourceScholar