← Search

Myungseo Song

3 accepted papers

2023

Unsupervised Pre-Training for Data-Efficient Text-to-Speech on Low Resource Languages

ICASSP 2023accepted

Neural text-to-speech (TTS) models can synthesize natural human speech when trained on large amounts of transcribed speech. How-ever, collecting such large-scale transcribed data is expensive. This paper proposes an unsupervised pre-training method for a sequence-to-sequence TTS model by leveraging…

Cited by 0SourceScholar
2021

Variable-Rate Deep Image Compression Through Spatially-Adaptive Feature Transform

ICCV 2021poster

We propose a versatile deep image compression network based on Spatial Feature Transform (SFT), which takes a source image and a corresponding quality map as inputs and produce a compressed image with variable rates. Our model covers a wide range of compression rates using a single model, which is c…

Cited by 119PDFcodeScholar