← Search

Michelle Tadmor Ramanovich

5 accepted papers

2025

SimulTron: On-Device Simultaneous Speech to Speech Translation

ICASSP 2025accepted

Simultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mobile devices remains a major challenge. We introduce SimulTron, a novel S2ST arch…

Cited by 0SourceScholar
2024

Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

ICLR 2024poster

We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire sy…

Cited by 40SourcePDFScholar
2024

Translatotron 3: Speech to Speech Translation with Monolingual Data

ICASSP 2024accepted

This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation. Experimental results in speech-to-speech translation tasks between Sp…

Cited by 0SourceScholar
2022

More Than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech

CVPR 2022poster

In this paper we present VDTTS, a Visually-Driven Text-to-Speech model. Motivated by dubbing, VDTTS takes advantage of video frames as an additional input alongside text, and generates speech that matches the video signal. We demonstrate how this allows VDTTS to, unlike plain TTS models, generate sp…

Cited by 21PDFScholar
2022

Translatotron 2: High-quality direct speech-to-speech translation with voice preservation

ICML 2022spotlight

We present Translatotron 2, a neural direct speech-to-speech translation model that can be trained end-to-end. Translatotron 2 consists of a speech encoder, a linguistic decoder, an acoustic synthesizer, and a single attention module that connects them together. Experimental results on three dataset…

Cited by 77SourcePDFScholar