← Search

Jonathan Le Roux

59 accepted papers

2025

Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos

ICASSP 2025accepted

Understanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstra…

Cited by 0SourceScholar
2025

Keeping the Balance: Anomaly Score Calculation for Domain Generalization

ICASSP 2025accepted

Emitted sounds may drastically change when using different microphones, when properties of the sound sources change, or when recording in different acoustic environments. Ideally, anomalous sound detection (ASD) systems should be able to generalize well to unseen target domains by only providing a f…

Cited by 0SourceScholar
2025

Leveraging Audio-Only Data for Text-Queried Target Sound Extraction

ICASSP 2025accepted

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text queries, the limited number of available high-quality text-…

Cited by 0SourceScholar
2025

No Class Left Behind: A Closer Look at Class Balancing for Audio Tagging

ICASSP 2025accepted

Large-scale audio tagging datasets like AudioSet usually suffer from severe class imbalance comprising many audio examples for common sound classes but only few examples of rare sound classes. The latter, however, may yet be equally or even more important to recognize. Therefore, it is common practi…

Cited by 0SourceScholar
2025

O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization

ICASSP 2025accepted

We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we develop a novel centroid refinement decoder whose usefulness is assessed through a rigorous ablation study. Our system p…

Cited by 0SourceScholar
2025

Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization

ICASSP 2025accepted

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields, spatial upsampling only from a few measured directions, e.g.…

Cited by 0SourceScholar
2025

Task-Aware Unified Source Separation

ICASSP 2025accepted

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or cinematic audio source separation (CASS) with a single model. These models are trained on large-scale data including spe…

Cited by 0SourceScholar
2024

Disentangled Acoustic Fields For Multimodal Physical Scene Understanding

IROS 2024poster

We study the problem of multimodal physical scene understanding, where an embodied agent needs to find fallen objects by inferring object properties, direction, and distance of an impact sound source. Previous works adopt feed-forward neural networks to directly regress the variables from sound, lea…

Cited by 0SourceScholar
2024

GLA-GRAD: A Griffin-Lim Extended Waveform Generation Diffusion Model

ICASSP 2024accepted

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a diffusion process for the generation of high-fidelity audio. How…

Cited by 0SourceScholar
2024

Generation or Replication: Auscultating Audio Latent Diffusion Models

ICASSP 2024accepted

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt at understanding the inner workings of audio latent diffusi…

Cited by 0SourceScholar
2024

Improving Audio Captioning Models with Fine-Grained Audio Features, Text Embedding Supervision, and LLM Mix-Up Augmentation

ICASSP 2024accepted

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a sequence-to-sequence (seq2seq) backbone powered by strong mo…

Cited by 0SourceScholar
2024

NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization

ICASSP 2024accepted

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have gained attention. Existing NF-based methods focused on estimat…

Cited by 0SourceScholar
2024

NeuroHeed+: Improving Neuro-Steered Speaker Extraction with Joint Auditory Attention Detection

ICASSP 2024accepted

Neuro-steered speaker extraction aims to extract the listener’s brainattended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recorded using electroencephalography (EEG) devices. Though promising, current methods…

Cited by 0SourceScholar
2024

RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual Navigation

CVPR 2024poster

We leverage Large Language Models (LLM) for zeroshot Semantic Audio Visual Navigation (SAVN). Existing methods utilize extensive training demonstrations for reinforcement learning yet achieve relatively low success rates and lack generalizability. The intermittent nature of auditory signals further…

Cited by 7SourcePDFScholar
2024

SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis

ICASSP 2024accepted

Generative adversarial network (GAN) models can synthesize high-quality audio signals while ensuring fast sample generation. However, they are difficult to train and are prone to several issues including mode collapse and divergence. In this paper, we introduce SpecDiff-GAN, a neural vocoder based o…

Cited by 0SourceScholar
2024

WI-FI based Indoor Monitoring Enhanced by Multimodal Fusion

ICASSP 2024accepted

Indoor monitoring systems are in high demand to protect vulnerable people, especially when they are alone at home, in nursing homes, hospitals, etc. Although surveillance systems in public spaces use cameras and microphones to find incidents, indoor monitoring in personal spaces needs to protect pri…

Cited by 0SourceScholar
2023

Latent Iterative Refinement for Modular Source Separation

ICASSP 2023accepted

Traditional source separation approaches train deep neural network models end-to-end with all the data available at once by minimizing the empirical risk on the whole training set. On the inference side, after training the model, the user fetches a static computation graph and runs the full model on…

Cited by 0SourceScholar
2023

Optimal Condition Training for Target Source Separation

ICASSP 2023accepted

Recent research has shown remarkable performance in leveraging multiple extraneous conditional and non-mutually-exclusive semantic concepts for sound source separation, allowing the flexibility to extract a given target source based on multiple different queries. In this work, we propose a new optim…

Cited by 0SourceScholar
2023

Reverberation as Supervision For Speech Separation

ICASSP 2023accepted

This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separation required the synthesis of mixtures of mixtures or assumed the existence of a teacher model, making them difficult to…

Cited by 9SourceScholar
2022

(2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering

AAAI 2022technical

Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approaches often ignore the fact that videos are essentially sequences of 2D ``views'' of events happening in a 3D space, and tha…

Cited by 25SourcePDFScholar
2022

Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy

ICASSP 2022accepted

Pseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to enhance the performance of end-to-end automatic speech recognition (ASR). Our prior work proposed momentum pseudo-labelin…

Cited by 14SourceScholar
2022

Audio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher Learning

ICASSP 2022accepted

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges (DSTC7, DSTC8). In these challenges, the best-performing systems…

Cited by 0SourceScholar
2022

Extended Graph Temporal Classification for Multi-Speaker End-to-End ASR

ICASSP 2022accepted

Graph-based temporal classification (GTC), a generalized form of the connectionist temporal classification loss, was recently proposed to improve automatic speech recognition (ASR) systems using graph-based supervision. For example, GTC was first used to encode an N-best list of pseudo-label sequenc…

Cited by 0SourceScholar
2022

Locate This, Not that: Class-Conditioned Sound Event DOA Estimation

ICASSP 2022accepted

Existing systems for sound event localization and detection (SELD) typically operate by estimating a source location for all classes at every time instant. In this paper, we propose an alternative class-conditioned SELD model for situations where we may not be interested in localizing all classes al…

Cited by 0SourceScholar
2022

The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks

ICASSP 2022accepted

The cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research. Recent efforts have mainly focused on separating speech from noise, speech from speech, musical instruments from each other, or sound events fr…

Cited by 0SourceScholar
2021

Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers

AAAI 2021technical

Given an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a human about the audio-visual content. This task thus poses a challenging multi-modal representation learning and reasoning…

Cited by 50SourcePDFScholar
2021

Semi-Supervised Speech Recognition Via Graph-Based Temporal Classification

ICASSP 2021accepted

Semi-supervised learning has demonstrated promising results in automatic speech recognition (ASR) by self-training using a seed ASR model with pseudo-labels generated for unlabeled data. The effectiveness of this approach largely relies on the pseudo-label accuracy, for which typically only the 1-be…

Cited by 0SourceScholar
2021

Transcription Is All You Need: Learning To Separate Musical Mixtures With Score As Supervision

ICASSP 2021accepted

Most music source separation systems require large collections of isolated sources for training, which can be difficult to obtain. In this work, we use musical scores, which are comparatively easy to obtain, as a weak label for training a source separation system. In contrast with previous score-inf…

Cited by 0SourceScholar
2021

Unsupervised Domain Adaptation for Speech Recognition via Uncertainty Driven Self-Training

ICASSP 2021accepted

The performance of automatic speech recognition (ASR) systems typically degrades significantly when the training and test data domains are mismatched. In this paper, we show that self-training (ST) combined with an uncertainty-based pseudo-label filtering approach can be effectively used for domain…

Cited by 0SourceScholar
2021

Visual Scene Graphs for Audio Source Separation

ICCV 2021poster

State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches often ignore the visual context of these sound sources or avoid modeling object interactions that may be useful to better ch…

Cited by 42PDFcodeScholar
2020

End-To-End Multi-Speaker Speech Recognition With Transformer

ICASSP 2020accepted

Recently, fully recurrent neural network (RNN) based end-to-end models have been proven to be effective for multi-speaker speech recognition in both the single-channel and multi-channel scenarios. In this work, we explore the use of Transformer models for these tasks by focusing on two aspects. Firs…

Cited by 0SourceScholar
2020

Unsupervised Speaker Adaptation Using Attention-Based Speaker Memory for End-to-End ASR

ICASSP 2020accepted

We propose an unsupervised speaker adaptation method inspired by the neural Turing machine for end-to-end (E2E) automatic speech recognition (ASR). The proposed model contains a memory block that holds speaker i-vectors extracted from the training data and reads relevant i-vectors from the memory th…

Cited by 0SourceScholar
2020

WHAMR!: Noisy and Reverberant Single-Channel Speech Separation

ICASSP 2020accepted

While significant advances have been made with respect to the separation of overlapping speech signals, studies have been largely constrained to mixtures of clean, near anechoic speech, not representative of many real-world scenarios. Although the WHAM! dataset introduced noise to the ubiquitous wsj…

Cited by 0SourceScholar
2019

Bootstrapping Single-channel Source Separation via Unsupervised Spatial Clustering on Stereo Mixtures

ICASSP 2019accepted

Separating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems based on deep learning are currently the most successful approaches for solving the underdetermined separation problem, w…

Cited by 0SourceScholar
2019

Class-conditional Embeddings for Music Source Separation

ICASSP 2019accepted

Isolating individual instruments in a musical mixture has a myriad of potential applications, and seems imminently achievable given the levels of performance reached by recent deep learning methods. While most musical source separation techniques learn an independent model for each instrument, we pr…

Cited by 0SourceScholar
2019

Cycle-consistency Training for End-to-end Speech Recognition

ICASSP 2019accepted

This paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionaries to build ASR systems, it still requires a large amount of paired data, i.e.,…

Cited by 0SourceScholar
2019

Teacher-student Deep Clustering for Low-delay Single Channel Speech Separation

ICASSP 2019accepted

The recently-proposed deep clustering algorithm introduced significant advances in monaural speaker-independent multi-speaker speech separation. Deep clustering operates on magnitude spectro-grams using bidirectional recurrent networks and K-means clustering, both of which require offline operation,…

Cited by 0SourceScholar
2019

The Phasebook: Building Complex Masks via Discrete Representations for Source Separation

ICASSP 2019accepted

Deep learning based speech enhancement and source separation systems have recently reached unprecedented levels of quality, to the point that performance is reaching a new ceiling. Most systems rely on estimating the magnitude of a target source, either directly or by computing a real-valued mask to…

Cited by 0SourceScholar
2018

An End-to-End Language-Tracking Speech Recognizer for Mixed-Language Speech

ICASSP 2018accepted

End-to-end automatic speech recognition (ASR) can significantly reduce the burden of developing ASR systems for new languages, by eliminating the need for linguistic information such as pronunciation dictionaries. This also creates an opportunity to build a monolithic multilingual ASR system with a…

Cited by 0SourceScholar
2018

End-to-End Multi-Speaker Speech Recognition

ICASSP 2018accepted

Current advances in deep learning have resulted in a convergence of methods across a wide range of tasks, opening the door for tighter integration of modules that were previously developed and optimized in isolation. Recent ground-breaking works have produced end-to-end deep network methods for both…

Cited by 0SourceScholar
2018

Multi-Channel Deep Clustering: Discriminative Spectral and Spatial Embeddings for Speaker-Independent Speech Separation

ICASSP 2018accepted

The recently-proposed deep clustering algorithm represents a fundamental advance towards solving the cocktail party problem in the single-channel case. When multiple microphones are available, spatial information can be leveraged to differentiate signals from different directions. This study combine…

Cited by 0SourceScholar
2017

BLSTM-HMM hybrid system combined with sound activity detection network for polyphonic Sound Event Detection

ICASSP 2017accepted

This paper presents a new hybrid approach for polyphonic Sound Event Detection (SED) which incorporates a temporal structure modeling technique based on a hidden Markov model (HMM) with a frame-by-frame detection method based on a bidirectional long short-term memory (BLSTM) recurrent neural network…

Cited by 0SourceScholar
2017

Deep clustering and conventional networks for music separation: Stronger together

ICASSP 2017accepted

Deep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging si…

Cited by 0SourceScholar
2017

Student-teacher network learning with enhanced features

ICASSP 2017accepted

Recent advances in distant-talking ASR research have confirmed that speech enhancement is an essential technique for improving the ASR performance, especially in the multichannel scenario. However, speech enhancement inevitably distorts speech signals, which can cause significant degradation when en…

Cited by 0SourceScholar
2016

Deep clustering: Discriminative embeddings for segmentation and separation

ICASSP 2016accepted

We address the problem of "cocktail-party" source separation in a deep learning framework called deep clustering. Previous deep network approaches to separation have shown promising performance in scenarios with a fixed number of sources, each belonging to a distinct signal class, such as speech and…

Cited by 0SourceScholar
2016

Full-Capacity Unitary Recurrent Neural Networks

NeurIPS 2016poster

Recurrent neural networks are powerful models for processing sequential data, but they are generally plagued by vanishing and exploding gradient problems. Unitary recurrent neural networks (uRNNs), which use unitary recurrence matrices, have recently been proposed as a means to avoid these issues. H…

2015

Micbots: Collecting large realistic datasets for speech and audio research using mobile robots

ICASSP 2015accepted

Speech and audio signal processing research is a tale of data collection efforts and evaluation campaigns. Large benchmark datasets for automatic speech recognition (ASR) have been instrumental in the advancement of speech recognition technologies. However, when it comes to robust ASR, source separa…

Cited by 0SourceScholar
2015

Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks

ICASSP 2015accepted

Separation of speech embedded in non-stationary interference is a challenging problem that has recently seen dramatic improvements using deep network-based methods. Previous work has shown that estimating a masking function to be applied to the noisy spectrum is a viable approach that can be improve…

Cited by 0SourceScholar