ICASSP 2018accepted0 citations
Fftnet: A Real-Time Speaker-Dependent Neural Vocoder
Zeyu Jin, Adam Finkelstein, Gautham J. Mysore, Jingwan Lu
Abstract
We introduce FFTNet, a deep learning approach synthesizing audio waveforms. Our approach builds on the recent WaveNet project, which showed that it was possible to synthesize a natural sounding audio waveform directly from a deep convolutional neural network. FFTNet offers two improvements over WaveNet. First it is substantially faster, allowing for real-time synthesis of audio waveforms. Second, when used as a vocoder, the resulting speech sounds more natural, as measured via a “mean opinion score” test.
BibTeX
@inproceedings{icassp2018_fftnetarealtimes,
title = {Fftnet: A Real-Time Speaker-Dependent Neural Vocoder},
author = {Zeyu Jin and Adam Finkelstein and Gautham J. Mysore and Jingwan Lu},
booktitle = {ICASSP 2018},
year = {2018}
}