ICASSP 2018accepted0 citations

Samplernn-Based Neural Vocoder for Statistical Parametric Speech Synthesis

Yang Ai, Hong-Chuan Wu, Zhen-Hua Ling

Abstract

This paper presents a SampleRNN-based neural vocoder for statistical parametric speech synthesis. This method utilizes a conditional SampleRNN model composed of a hierarchical structure of GRU layers and feed-forward layers to capture long-span dependencies between acoustic features and waveform sequences. Compared with conventional vocoders based on the source-filter model, our proposed vocoder is trained without assumptions derived from the prior knowledge of speech production and is able to provide a better modeling and recovery of phase information. Objective and subjective evaluations are conducted on two corpora. Experimental results suggested that our proposed vocoder can achieve higher quality of synthetic speech than the STRAIGHT vocoder and a WaveNet-based neural vocoder with similar run-time efficiency, no matter natural or predicted acoustic features are used as inputs.

BibTeX
@inproceedings{icassp2018_samplernnbasedne,
  title = {Samplernn-Based Neural Vocoder for Statistical Parametric Speech Synthesis},
  author = {Yang Ai and Hong-Chuan Wu and Zhen-Hua Ling},
  booktitle = {ICASSP 2018},
  year = {2018}
}
Samplernn-Based Neural Vocoder for Statistical Parametric Speech Synthesis · ICASSP 2018