← Search

Bajibabu Bollepalli

6 accepted papers

2022

Distribution Augmentation for Low-Resource Expressive Text-To-Speech

ICASSP 2022accepted

This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data. Our goal is to in-crease diversity of text conditionings available during training. This helps to reduce overfitting, e…

Cited by 0SourceScholar
2019

Waveform Generation for Text-to-speech Synthesis Using Pitch-synchronous Multi-scale Generative Adversarial Networks

ICASSP 2019accepted

The state-of-the-art in text-to-speech (TTS) synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference process, while their parallel versions are difficult to train and even more c…

Cited by 24SourceScholar
2018

Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks

ICASSP 2018accepted

This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from M…

Cited by 0SourceScholar
2017

Frequency-warped time-weighted linear prediction for glottal vocoding

ICASSP 2017accepted

Auto-regressive modeling is a prevalent source-filter separation method of speech. Conventional linear prediction (LP) and its derivatives such as weighted linear prediction (WeLP) produce parametric spectral models within a linear frequency scale, whereas frequency-warped linear prediction (WaLP) c…

Cited by 0SourceScholar
2017

Lombard speech synthesis using long short-term memory recurrent neural networks

ICASSP 2017accepted

In statistical parametric speech synthesis (SPSS), a few studies have investigated the Lombard effect, specifically by using hidden Markov model (HMM)-based systems. Recently, artificial neural networks have demonstrated promising results in SPSS, specifically by using long short-term memory recurre…

Cited by 0SourceScholar
2016

High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network

ICASSP 2016accepted

Achieving high quality and naturalness in statistical parametric synthesis of female voices remains to be difficult despite recent advances in the study area. Vocoding is one such key element in all statistical speech synthesizers that is known to affect the synthesis quality and naturalness. The pr…

Cited by 0SourceScholar