Autovocoder: Fast Waveform Generation from a Learned Speech Representation Using Differentiable Digital Signal Processing
Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation.A mel-spectrogram is extracted from the waveform by a simple, fast DSP operation, but generating a high-quality waveform from…