← Search

Prashant Serai

5 accepted papers

2024

Ultra-Lightweight Neural Differential DSP Vocoder for High Quality Speech Synthesis

ICASSP 2024accepted

Neural vocoders model the raw audio waveform and synthesize high-quality audio, but even the highly efficient ones, like MB-MelGAN and LPCNet, fail to run real-time on a low-end device like a smartglass. A pure digital signal processing (DSP) based vocoder can be implemented via lightweight fast Fou…

Cited by 5SourceScholar
2023

Voice-Preserving Zero-Shot Multiple Accent Conversion

ICASSP 2023accepted

Most people who have tried to learn a foreign language would have experienced difficulties understanding or speaking with a native speaker’s accent. For native speakers, understanding or speaking a new accent is likewise a difficult task. An accent conversion system that changes a speaker’s accent b…

Cited by 26SourceScholar
2020

End to End Speech Recognition Error Prediction with Sequence to Sequence Learning

ICASSP 2020accepted

Simulating the errors made by a speech recognizer on plain text has proven useful to help train downstream NLP tasks to be robust to real ASR errors at test time. Prior work in this domain has focused on modeling confusions at the phonetic level, and using a lexicon to convert from words to phones a…

Cited by 0SourceScholar
2019

Improving Human-computer Interaction in Low-resource Settings with Text-to-phonetic Data Augmentation

ICASSP 2019accepted

Off-the-shelf speech recognition systems can yield useful results and accelerate application development, but general-purpose systems applied to specialized domains can introduce acoustically small-but semantically catastrophic-errors. Furthermore, sufficient audio data may not be available to devel…

Cited by 8SourceScholar
2019

Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers

ICASSP 2019accepted

Modeling the errors of a speech recognizer can help simulate errorful recognized speech data from plain text, which has proven useful for tasks like discriminative language modeling, improving robustness of NLP systems, where limited or even no audio data is available at train time. Previous work ty…

Cited by 0SourceScholar