Improving Recognition-Synthesis Based any-to-one Voice Conversion with Cyclic Training
In recognition-synthesis based any-to-one voice conversion (VC), an automatic speech recognition (ASR) model is employed to extract content-related features and a synthesizer is built to predict the acoustic features of the target speaker from the content-related features of any source speakers at t…