← Search

June Sig Sung

2 accepted papers

2023

Investigating Content-Aware Neural Text-to-Speech MOS Prediction Using Prosodic and Linguistic Features

ICASSP 2023accepted

Current state-of-the-art methods for automatic synthetic speech evaluation are based on MOS prediction neural models. Such MOS prediction models include MOSNet and LDNet that use spectral features as input, and SSL-MOS that relies on a pretrained selfsupervised learning model that directly uses the…

Cited by 0SourceScholar
2021

Prosodic Clustering for Phoneme-Level Prosody Control in End-to-End Speech Synthesis

ICASSP 2021accepted

This paper presents a method for controlling the prosody at the phoneme level in an autoregressive attention-based text-to-speech system. Instead of learning latent prosodic features with a variational framework as is commonly done, we directly extract phoneme-level F0 and duration features from the…

Cited by 0SourceScholar