← Search

Slava Shechtman

5 accepted papers

2025

A Non-autoregressive Model for Joint STT and TTS

ICASSP 2025accepted

In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text modalities as input either individually or together. The proposed mo…

Cited by 0SourceScholar
2024

Speak While You Think: Streaming Speech Synthesis During Text Generation

ICASSP 2024accepted

Large Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outputs typically results in notable latency, which is impractical for fluent voice conversations. We propose LLM2Speech, an…

Cited by 0SourceScholar
2021

Stable Checkpoint Selection and Evaluation in Sequence to Sequence Speech Synthesis

ICASSP 2021accepted

Autoregressive Attentive Sequence-to-Sequence (S2S) speech synthesis is considered state-of-the-art in terms of speech quality and naturalness, as evaluated on a finite set of testing utterances. However, it can occasionally suffer from stability issues at inference time, such as local intelligibili…

Cited by 0SourceScholar
2015

Coherent modification of pitch and energy for expressive prosody implantation

ICASSP 2015accepted

In expressive TTS and voice transformation systems, implantation of expressive prosody derived from external out-of-domain sources often leads to extreme pitch modification that compromises the naturalness of the synthesized speech. In this work we investigate and prove a hypothesis that the natural…

Cited by 0SourceScholar