← Search

Neil Zeghidour

24 accepted papers

2026

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

ICML 2026poster

Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguished by their real-time interactivity, including handling of pauses, interruptions, and backchannels. However, improving their factuality remains an …

Cited by 0SourceScholar
2026

Simultaneous Speech-to-Speech Translation Without Aligned Data

ICML 2026oral

Simultaneous speech translation is the task of translating source speech into a target language in real-time. Given that the dependencies between source and target words are non-monotonic (e.g. the word order can change between German and English), this means learning to jointly align and translate.…

Cited by 0SourcecodeScholar
2026

Vision-Speech Models: Teaching Speech Models to Converse about Images

CVPR 2026

The recent successes of Vision-Language models raise the question of how to equivalently imbue a pretrained speech model with vision understanding, an important milestone towards building a multimodal speech model able to freely converse about images. Building such a conversational Vision-Speech mod

Cited by 0SourcecodeScholar
2025

Aligning Spoken Dialogue Models from User Interactions

ICML 2025poster

We propose a novel preference alignment framework for improving spoken dialogue models on real-time conversations from user interactions. Current preference learning methods primarily focus on text-based language models, and are not directly suited to the complexities of real-time speech interaction…

Cited by 0SourcePDFScholar
2025

High-Fidelity Simultaneous Speech-To-Speech Translation

ICML 2025poster

We introduce Hibiki, a decoder-only model for simultaneous speech translation. Hibiki leverages a multistream language model to synchronously process source and target speech, and jointly produces text and audio tokens to perform speech-to-text and speech-to-speech translation. We furthermore addres…

2025

MAD Speech: Measures of Acoustic Diversity of Speech

NAACL 2025long

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse remains unclear due to a lack of appropriate metrics. We address…

2024

MusicRL: Aligning Music Generation to Human Preferences

ICML 2024poster

We propose MusicRL, the first music generation system finetuned from human feedback. Appreciation of text-to-music models is particularly subjective since the concept of musicality as well as the specific intention behind a caption are user-dependent (e.g. a caption such as “upbeat workout music” ca…

2023

Disentangling Speech from Surroundings with Neural Embeddings

ICASSP 2023accepted

We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec. We introduce a new training procedure that allows our model to produce structured encodings of audio waveforms given by embedding vectors, where one part of the embedding vector rep…

Cited by 19SourceScholar
2023

LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models

ICASSP 2023accepted

We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes audio into a hierarchy of coarse-to-fine tokens using residual vector quantization. LMCodec trains a Transformer language…

Cited by 0SourceScholar
2023

Pose-graph SLAM Using Multi-order Ultrasonic Echoes and Beamforming for Long-range Inspection Robots

ICRA 2023poster

This paper presents a Graph-based Simultaneous Localization And Mapping (GraphSLAM) approach for a robotic system relying on the reflections of ultrasonic guided waves to enable long-range inspection tasks on plate-based metal structures. A measurement model that can leverage multi-order acoustic ec…

Cited by 1SourceScholar
2023

Speech Intelligibility Classifiers from 550k Disordered Speech Samples

ICASSP 2023accepted

We developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a five-point scale. We trained three models following differen…

Cited by 0SourceScholar
2022

Combined Grid and Feature-based Mapping of Metal Structures with Ultrasonic Guided Waves

ICRA 2022poster

The ultrasonic mapping of plate-based facilities is an essential step towards the robotic inspection of large metal structures such as storage tanks or ship hulls. This work proposes a novel framework that exploits ultrasonic echoes to recover grid-based and feature-based spatial representations joi…

Cited by 2SourceScholar
2022

General-purpose, long-context autoregressive modeling with Perceiver AR

ICML 2022spotlight

Real-world data is high-dimensional: a book, image, or musical performance can easily contain hundreds of thousands of elements even after compression. However, the most commonly used autoregressive models, Transformers, are prohibitively expensive to scale to the number of inputs and layers needed…

2022

Learning Strides in Convolutional Neural Networks

ICLR 2022oral

Convolutional neural networks typically contain several downsampling operators, such as strided convolutions or pooling layers, that progressively reduce the resolution of intermediate representations. This provides some shift-invariance while reducing the computational complexity of the whole archi…

2021

LEAF: A Learnable Frontend for Audio Classification

ICLR 2021poster

Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by the fundamental limitations of handmade representations. In this work we show that…

2021

Learning From Heterogeneous Eeg Signals with Differentiable Channel Reordering

ICASSP 2021accepted

We propose CHARM, a method for training a single neural network across inconsistent input channels. Our work is motivated by Electroencephalography (EEG), where data collection protocols from different headsets result in varying channel ordering and number, which limits the feasibility of transferri…

Cited by 0SourceScholar
2019

To Reverse the Gradient or Not: an Empirical Comparison of Adversarial and Multi-task Learning in Speech Recognition

ICASSP 2019accepted

Transcribed datasets typically contain speaker identity for each instance in the data. We investigate two ways to incorporate this information during training: Multi-Task Learning and Adversarial Learning. In multi-task learning, the goal is speaker prediction; we expect a performance improvement wi…

Cited by 0SourceScholar
2018

Learning Filterbanks from Raw Speech for Phone Recognition

ICASSP 2018accepted

We train a bank of complex filters that operates on the raw waveform and is fed into a convolutional neural network for end-to-end phone recognition. These time-domain filterbanks (TD-filterbanks) are initialized as an approximation of mel-filterbanks, and then fine-tuned jointly with the remaining…

Cited by 0SourceScholar
2018

SING: Symbol-to-Instrument Neural Generator

NeurIPS 2018poster

Recent progress in deep learning for audio synthesis opens the way to models that directly produce the waveform, shifting away from the traditional paradigm of relying on vocoders or MIDI synthesizers for speech or music generation. Despite their successes, current state-of-the-art neural audio synt…

2017

Fader Networks:Manipulating Images by Sliding Attributes

NeurIPS 2017poster

This paper introduces a new encoder-decoder architecture that is trained to reconstruct images by disentangling the salient information of the image and the values of attributes directly in the latent space. As a result, after training, our model can generate different realistic versions of an input…

2016

A deep scattering spectrum - Deep Siamese network pipeline for unsupervised acoustic modeling

ICASSP 2016accepted

Recent work has explored deep architectures for learning acoustic features in an unsupervised or weakly-supervised way for phone recognition. Here we investigate the role of the input features, and in particular we test whether standard mel-scaled filterbanks could be replaced by inherently richer r…

Cited by 0SourceScholar