← Search

Vladimir Bataev

3 accepted papers

2025

HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR

ICLR 2025poster

We present Hybrid-Autoregressive INference TrANsducers (HAINAN), a novel architecture for speech recognition that extends the Token-and-Duration Transducer (TDT) model. Trained with randomly masked predictor network outputs, HAINAN supports both autoregressive inference with all network components a…

Cited by 0SourcePDFScholar
2025

TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer

ICASSP 2025accepted

This work introduces TTS-Transducer – a novel architecture for text-to-speech, leveraging the strengths of audio codec models and neural transducers. Transducers, renowned for their superior quality and robustness in speech recognition, are employed to learn monotonic alignments and allow for avoidi…

Cited by 0SourceScholar
2023

Powerful and Extensible WFST Framework for Rnn-Transducer Losses

ICASSP 2023accepted

This paper presents a framework based on Weighted Finite-State Transducers (WFST) to simplify the development of modifications for RNN-Transducer (RNN-T) loss. Existing implementations of RNN-T use CUDA-related code, which is hard to extend and debug. WFSTs are easy to construct and extend, and allo…

Cited by 0SourceScholar