← Search

Mehrez Souden

3 accepted papers

2025

ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model

ICASSP 2025accepted

We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sound objects. ImmerseDiffusion is trained to generate first-order ambisonics (FOA) audio, which is a conventional spatial…

Cited by 0SourceScholar
2024

Resource-Constrained Stereo Singing Voice Cancellation

ICASSP 2024accepted

We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to large state-of-the-art source separation networks starting from a small, efficient…

Cited by 0SourceScholar
2021

Dynamic Curriculum Learning via Data Parameters for Noise Robust Keyword Spotting

ICASSP 2021accepted

We propose dynamic curriculum learning via data parameters for noise robust keyword spotting. Data parameter learning has recently been introduced for image processing, where weight parameters, so-called data parameters, for target classes and instances are introduced and optimized along with model…

Cited by 0SourceScholar