← Search

Mounya Elhilali

12 accepted papers

2025

SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

ICASSP 2025accepted

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both a…

Cited by 0SourceScholar
2024

Biomimetic Mappings for Active Sonar Object Recognition in Clutter

ICASSP 2024accepted

SONAR technology plays a pivotal role in terrain exploration and specifically identification of objects of interest. However, it grapples with a recurring challenge of clutter and noise which limits the performance of target recognition models. The challenge of noisy observations renders the choice…

Cited by 0SourceScholar
2024

DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction

ICASSP 2024accepted

Common target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separating the target from the background. This study introduces DPM-TSE, a generative…

Cited by 0SourceScholar
2024

Investigating Self-Supervised Deep Representations for EEG-Based Auditory Attention Decoding

ICASSP 2024accepted

Auditory Attention Decoding (AAD) algorithms play a crucial role in isolating desired sound sources within challenging acoustic environments directly from brain activity. Although recent research has shown promise in AAD using shallow representations such as auditory envelope and spectrogram, there…

Cited by 0SourceScholar
2021

Self-Training for Sound Event Detection in Audio Mixtures

ICASSP 2021accepted

Sound event detection (SED) takes on the task of identifying presence of specific sound events in a complex audio recording. SED has tremendous implications in video analytics, smart speaker algorithms and audio tagging. Recent advances in deep learning have afforded remarkable advances in performan…

Cited by 0SourceScholar
2020

Synthesizing Engaging Music Using Dynamic Models of Statistical Surprisal

ICASSP 2020accepted

Synthesis of music content generally leverages the underlying statistical structure of music to develop generative models, able to create new musical expressions within the same genre. In this work, we explore the statistical structure of a musical corpus and its effect on modulating the attention o…

Cited by 0SourceScholar
2019

Joint Acoustic and Class Inference for Weakly Supervised Sound Event Detection

ICASSP 2019accepted

Sound event detection is a challenging task, especially for scenes with multiple simultaneous events. While event classification methods tend to be fairly accurate, event localization presents additional challenges, especially when large amounts of labeled data are not available. Task4 of the 2018 D…

Cited by 0SourceScholar