← Search

Vitaly Lavrukhin

10 accepted papers

2025

Anticipating Future with Large Language Model for Simultaneous Machine Translation

NAACL 2025long

Simultaneous machine translation (SMT) takes streaming input utterances and incrementally produces target text. Existing SMT methods only use the partial utterance that has already arrived at the input and the generated hypothesis. Motivated by human interpreters’ technique to forecast future words…

Cited by 0SourcePDFScholar
2025

Chain-of-Thought Prompting for Speech Translation

ICASSP 2025accepted

Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddings for prompting, resulting in Speech-LLM models that exhibit strong performance…

Cited by 0SourceScholar
2025

EMMeTT: Efficient Multimodal Machine Translation Training

ICASSP 2025accepted

A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and proposes a joint multimodal training regime of Speech-LLM to include automatic s…

Cited by 0SourceScholar
2025

Extending Automatic Machine Translation Evaluation to Book-Length Documents

EMNLP 2025

Despite Large Language Models (LLMs) demonstrating superior translation performance and long-context capabilities, evaluation methodologies remain constrained to sentence-level assessment due to dataset limitations, token number restrictions in metrics, and rigid sentence boundary requirements. We i

2025

Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic

ICASSP 2025accepted

Despite Arabic being one of the most widely spoken languages, the development of Arabic Automatic Speech Recognition (ASR) systems faces significant challenges due to the language’s complexity, and only a limited number of public Arabic ASR models exist. While much of the focus has been on Modern St…

Cited by 0SourceScholar
2025

TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer

ICASSP 2025accepted

This work introduces TTS-Transducer – a novel architecture for text-to-speech, leveraging the strengths of audio codec models and neural transducers. Transducers, renowned for their superior quality and robustness in speech recognition, are employed to learn monotonic alignments and allow for avoidi…

Cited by 0SourceScholar
2024

A Chat about Boring Problems: Studying GPT-Based Text Normalization

ICASSP 2024accepted

Text normalization - the conversion of text from written to spoken form - is traditionally assumed to be an ill-formed task for language modeling. In this work, we argue otherwise. We empirically show the capacity of Large-Language Models (LLM) for text normalization in few-shot scenarios. Combining…

Cited by 0SourceScholar
2023

Conformer-Based Target-Speaker Automatic Speech Recognition For Single-Channel Audio

ICASSP 2023accepted

We propose CONF-TSASR, a non-autoregressive end-to-end time-frequency domain architecture for single-channel target-speaker automatic speech recognition (TS-ASR). The model consists of a TitaNet based speaker embedding module, a Conformer based masking as well as ASR modules. These modules are joint…

Cited by 0SourceScholar
2020

Quartznet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions

ICASSP 2020accepted

We propose a new end-to-end neural acoustic model for automatic speech recognition. The model is composed of multiple blocks with residual connections between them. Each block consists of one or more modules with 1D time-channel separable convolutional layers, batch normalization, and ReLU layers. I…

Cited by 0SourceScholar