ICASSP 2023accepted0 citations

Performance Comparison of TTS Models for Brazilian Portuguese to Establish a Baseline

Wilmer Lobato, Felipe Farias, William Cruz, Marcellus Amadeus

Abstract

This paper compares the performance of three text-to-speech (TTS) models released from June 2021 to January 2022 in order to establish a baseline for Brazilian Portuguese. Those models were trained using dataset for Brazilian Portuguese. The experimental setup considers tts-portuguese dataset to fine-tune the following TTS models: VITS end-to-end model; glowtts and gradtts acoustic models both using hifigan vocoder. Performance metrics are arranged into objective and subjective metrics. As subjective metrics, the naturalness and intelligibility are measured based on the mean opinion score (MOS). Results shows that gradtts+hifigan model achieved naturalness of 4.07 MOS, close to performance of current commercial models.

BibTeX
@inproceedings{icassp2023_performancecompa,
  title = {Performance Comparison of TTS Models for Brazilian Portuguese to Establish a Baseline},
  author = {Wilmer Lobato and Felipe Farias and William Cruz and Marcellus Amadeus},
  booktitle = {ICASSP 2023},
  year = {2023}
}