← Search

Stanislav Beliaev

2 accepted papers

2022

Mixer-TTS: Non-Autoregressive, Fast and Compact Text-to-Speech Model Conditioned on Language Model Embeddings

ICASSP 2022accepted

This paper describes Mixer-TTS, a non-autoregressive model for mel-spectrogram generation. The model is based on the MLP-Mixer architecture adapted for speech synthesis. The basic Mixer-TTS contains pitch and duration predictors, with the latter being trained with an unsupervised TTS alignment frame…

Cited by 18SourceScholar
2020

Quartznet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions

ICASSP 2020accepted

We propose a new end-to-end neural acoustic model for automatic speech recognition. The model is composed of multiple blocks with residual connections between them. Each block consists of one or more modules with 1D time-channel separable convolutional layers, batch normalization, and ReLU layers. I…

Cited by 331SourceScholar