← Search

Eliya Nachmani

20 accepted papers

2026

Token-Based Audio Inpainting via Discrete Diffusion

ICLR 2026poster

Audio inpainting seeks to restore missing segments in degraded recordings. Previous diffusion-based methods exhibit impaired performance when the missing region is large. We introduce the first approach that applies discrete diffusion over tokenized music representations from a pre-trained audio tok…

Cited by 0SourceScholar
2025

SimulTron: On-Device Simultaneous Speech to Speech Translation

ICASSP 2025accepted

Simultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mobile devices remains a major challenge. We introduce SimulTron, a novel S2ST arch…

Cited by 0SourceScholar
2024

Separate and Diffuse: Using a Pretrained Diffusion Model for Better Source Separation

ICLR 2024poster

The problem of speech separation, also known as the cocktail party problem, refers to the task of isolating a single speech signal from a mixture of speech signals. Previous work on source separation derived an upper bound for the source separation task in the domain of human speech. This bound is d…

Cited by 7SourcePDFScholar
2024

Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

ICLR 2024poster

We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire sy…

Cited by 40SourcePDFScholar
2024

Translatotron 3: Speech to Speech Translation with Monolingual Data

ICASSP 2024accepted

This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation. Experimental results in speech-to-speech translation tasks between Sp…

Cited by 0SourceScholar
2023

Decision S4: Efficient Sequence-Based RL via State Spaces Layers

ICLR 2023poster

Recently, sequence learning methods have been applied to the problem of off-policy Reinforcement Learning, including the seminal work on Decision Transformers, which employs transformers for this task. Since transformers are parameter-heavy, cannot benefit from history longer than a fixed window siz…

Cited by 30SourcePDFScholar
2023

kNN-Diffusion: Image Generation via Large-Scale Retrieval

ICLR 2023poster

Recent text-to-image models have achieved impressive results. However, since they require large-scale datasets of text-image pairs, it is impractical to train them on new domains where data is scarce or not labeled. In this work, we propose using large-scale retrieval methods, in particular, efficie…

Cited by 139SourcePDFScholar
2021

Single Channel Voice Separation for Unknown Number of Speakers Under Reverberant and Noisy Settings

ICASSP 2021accepted

We present a unified network for voice separation of an unknown number of speakers. The proposed approach is composed of several separation heads optimized together with a speaker classification branch. The separation is carried out in the time domain, together with parameter sharing between all sep…

Cited by 0SourceScholar
2018

VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop

ICLR 2018poster

We present a new neural text to speech (TTS) method that is able to transform text to speech in voices that are sampled in the wild. Unlike other systems, our solution is able to deal with unconstrained voice samples and without requiring aligned phonemes or linguistic features. The network architec…