← Search

Lianhong Cai

12 accepted papers

2018

Applying Multitask Learning to Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English Speech

ICASSP 2018accepted

For mispronunciation detection and diagnosis (MDD), nowadays approaches generally treat the phonemes in correct and mispronunciations as the same despite the fact they may actually carry different characteristics. Furthermore, serious data imbalance issue between correct and mispronunciation in data…

Cited by 0SourceScholar
2018

Emphatic Speech Generation with Conditioned Input Layer and Bidirectional LSTMS for Expressive Speech Synthesis

ICASSP 2018accepted

By highlighting the focus of an utterance to draw attention, emphasis in speech interaction plays an important role for speaker intention expressing and understanding. Therefore, emphatic speech synthesis draws increasing interest in the text-to-speech (TTS) area. For emphatic speech synthesis, thre…

Cited by 0SourceScholar
2017

Learning cross-lingual knowledge with multilingual BLSTM for emphasis detection with limited training data

ICASSP 2017accepted

Bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) has achieved state-of-the-art performance in many sequence processing problems given its capability in capturing contextual information. However, for languages with limited amount of training data, it is still difficult to o…

Cited by 0SourceScholar
2017

Multi-task learning of structured output layer bidirectional LSTMS for speech synthesis

ICASSP 2017accepted

Recurrent neural networks (RNNs) and their bidirectional long short term memory (BLSTM) variants are powerful sequence modelling approaches. Their inherently strong ability in capturing long range temporal dependencies allow BLSTM-RNN speech synthesis systems to produce higher quality and smoother s…

Cited by 0SourceScholar
2016

A deep bidirectional long short-term memory based multi-scale approach for music dynamic emotion prediction

ICASSP 2016accepted

Music Dynamic Emotion Prediction is a challenging and significant task. In this paper, We adopt the dimensional valence-arousal (V-A) emotion model to represent the dynamic emotion in music. Considering the high context correlation among the music feature sequence and the advantage of Bidirectional…

Cited by 0SourceScholar
2016

Learning cross-lingual information with multilingual BLSTM for speech synthesis of low-resource languages

ICASSP 2016accepted

Bidirectional long short-term memory (BLSTM) based speech synthesis has shown great potential in improving the quality of the synthetic speech. However, for low-resource languages, it is difficult to obtain a high quality BLSTM model. BLSTM based speech synthesis can be viewed as a transformation be…

Cited by 0SourceScholar
2016

Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatar

ICASSP 2016accepted

Speech is bimodal in nature. There are close correlations between the acoustic speech signals and the visual gestures such as lip movements, facial expressions and head motions. For speech driven talking avatar, how to derive more representative acoustic features from which to predict more accurate…

Cited by 0SourceScholar
2016

Question detection from acoustic features using recurrent neural network with gated recurrent unit

ICASSP 2016accepted

Question detection is of importance for many speech applications. Only parts of the speech utterances can provide useful clues for question detection. Previous work of question detection using acoustic features in Mandarin conversation is weak in capturing such proper time context information, which…

Cited by 0SourceScholar
2016

SVR based double-scale regression for dynamic emotion prediction in music

ICASSP 2016accepted

Dynamic music emotion prediction is to recognize the continuous emotion contained in music, and has various applications. In recent years, dynamic music emotion recognition is widely studied, while the inside structure of the emotion in music remains unclear. We conduct a data observation based on t…

Cited by 0SourceScholar
2015

A deep recurrent approach for acoustic-to-articulatory inversion

ICASSP 2015accepted

To solve the acoustic-to-articulatory inversion problem, this paper proposes a deep bidirectional long short term memory recurrent neural network and a deep recurrent mixture density network. The articulatory parameters of the current frame may have correlations with the acoustic features many frame…

Cited by 0SourceScholar
2015

HMM-based emphatic speech synthesis for corrective feedback in computer-aided pronunciation training

ICASSP 2015accepted

This paper investigates the incorporation of hidden Markov model (HMM) based emphatic speech synthesis for audio exaggeration into an audio-visual speech synthesis framework for the corrective feedback in computer-aided pronunciation training (CAPT). To improve the voice quality of the synthetic emp…

Cited by 0SourceScholar