← Search

Frank K. Soong

21 accepted papers

2022

A Universal Ordinal Regression for Assessing Phoneme-Level Pronunciation

ICASSP 2022accepted

The efficacy and robustness of Ordinal Regression (OR) in assessing speech pronunciation for language learning at phrase level has been shown before. However, for assessing phoneme pronunciation, we need to: 1. collect human scoring annotations for phoneme tokens of a short duration (60-70 ms); 2. t…

Cited by 8SourceScholar
2022

An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) Embeddings

ICASSP 2022accepted

Many mispronunciation detection and diagnosis (MD&D) research approaches try to exploit both the acoustic and linguistic features as input. Yet the improvement of the performance is limited, partially due to the shortage of large amount annotated training data at the phoneme level. Phonetic embeddin…

Cited by 0SourceScholar
2022

Improving Fastspeech TTS with Efficient Self-Attention and Compact Feed-Forward Network

ICASSP 2022accepted

FastSpeech, as a feed-forward transformer based TTS, can avoid the slow serial, autoregressive inference to generate the target mel-spectrogram in a parallel way. As a non-autoregressive TTS, the latency and computation load in inference is shifted from vocoder to transformer where the efficiency is…

Cited by 0SourceScholar
2021

A New High Quality Trajectory Tiling Based Hybrid TTS In Real Time

ICASSP 2021accepted

A trajectory tiling based, hybrid TTS is revisited in this study for improving its synthesis performance. A combination of Transformer encoder and RNN based decoder architecture where two-level, at both word and Chinese phonetic alphabet letter levels, linguistic representation is exploited to gener…

Cited by 0SourceScholar
2021

Improving Pronunciation Assessment Via Ordinal Regression with Anchored Reference Samples

ICASSP 2021accepted

Sentence level pronunciation assessment is important for Computer Assisted Language Learning (CALL). Traditional speech pronunciation assessment, based on the Goodness of Pronunciation (GOP) algorithm, has some weakness in assessing a speech utterance: 1) Phoneme GOP scores cannot be easily translat…

Cited by 0SourceScholar
2021

MBNET: MOS Prediction for Synthesized Speech with Mean-Bias Network

ICASSP 2021accepted

Mean opinion score (MOS) is a popular subjective metric to assess the quality of synthesized speech, and usually involves multiple human judges to evaluate each speech utterance. To reduce the labor cost in MOS test, multiple methods have been proposed to automatically predict MOS scores. To our kno…

Cited by 0SourceScholar
2020

An Improved Frame-Unit-Selection Based Voice Conversion System Without Parallel Training Data

ICASSP 2020accepted

A frame-unit-selection based voice conversion system proposed earlier by us is revisited here to enhance its performance in both speech naturalness and speaker similarity. Speaker independent, bilingual (Mandarin Chinese and American English) deep neural net (DNN) acoustic model’s output, frame-leve…

Cited by 0SourceScholar
2020

Improving LPCNET-Based Text-to-Speech with Linear Prediction-Structured Mixture Density Network

ICASSP 2020accepted

In this paper, we propose an improved LPCNet vocoder using a linear prediction (LP)-structured mixture density network (MDN). The recently proposed LPCNet vocoder has successfully achieved high-quality and lightweight speech synthesis systems by combining a vocal tract LP filter with a WaveRNN-based…

Cited by 0SourceScholar
2020

Improving Prosody with Linguistic and Bert Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTS

ICASSP 2020accepted

Recent advances of neural TTS have made "human parity" synthesized speech possible when a large amount of studio-quality training data from a voice talent is available. However, with only limited, casual recordings from an ordinary speaker, human-like TTS is still a big challenge, in addition to oth…

Cited by 0SourceScholar
2019

Domain Adversarial Training for Improving Keyword Spotting Performance of ESL Speech

ICASSP 2019accepted

A second language (L2) learner usually cannot speak L2 well in both pronunciations and forming-of-words. Hence his/her L2 speech cannot be well recognized by a recognizer trained with native data. Domain adversarial training (DAT), capable of reducing the acoustic mismatch between training and testi…

Cited by 0SourceScholar
2019

NN-based Ordinal Regression for Assessing Fluency of ESL Speech

ICASSP 2019accepted

Automatic assessment of a language learner's speech fluency is highly desirable for language education, e.g. for English as a Second Language (ESL) learning. In this paper, we formulate the fluency assessment as a problem of Ordinal Regression with Anchored Reference Samples (ORARS), where the fluen…

Cited by 0SourceScholar
2018

Exploring Sequential Characteristics in Speaker Bottleneck Feature for Text-Dependent Speaker Verification

ICASSP 2018accepted

In this paper, given the speaker bottleneck feature vectors extracted with speaker discriminant neural networks, we focus on using the sequential speaker characteristics for text-dependent speaker verification. In each evaluation trial, speaker supervectors are used as the representations of the seq…

Cited by 0SourceScholar
2015

AA spectral space warping approach to cross-lingual voice transformation in HMM-based TTS

ICASSP 2015accepted

This paper presents a new approach to cross-lingual voice transformation in HMM-based TTS with only the recordings from two monolingual speakers in different languages (e.g. Mandarin and English). We aim to synthesize one speaker's speech in the other language. We regard the spectral space of any sp…

Cited by 0SourceScholar
2015

Multi-speaker modeling and speaker adaptation for DNN-based TTS synthesis

ICASSP 2015accepted

In DNN-based TTS synthesis, DNNs hidden layers can be viewed as deep transformation for linguistic features and the output layers as representation of acoustic space to regress the transformed linguistic features to acoustic parameters. The deep-layered architectures of DNN can not only represent hi…

Cited by 0SourceScholar
2015

Word embedding for recurrent neural network based TTS synthesis

ICASSP 2015accepted

The current state of the art TTS synthesis can produce synthesized speech with highly decent quality if rich segmental and suprasegmental information are given. However, some suprasegmental features, e.g., Tone and Break (TOBI), are time consuming due to being manually labeled with a high inconsiste…

Cited by 0SourceScholar