← Search

Mostafa Sadeghi

11 accepted papers

2025

Diffusion-based Unsupervised Audio-visual Speech Enhancement

ICASSP 2025accepted

This paper proposes a new unsupervised audiovisual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion model is pre-trained on clean speech conditioned on correspondin…

Cited by 12SourceScholar
2024

A Weighted-Variance Variational Autoencoder Model for Speech Enhancement

ICASSP 2024accepted

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed for the generative model, where the speech information is encoded in the varianc…

Cited by 0SourceScholar
2024

Diffusion-Based Speech Enhancement with a Weighted Generative-Supervised Learning Loss

ICASSP 2024accepted

Diffusion-based generative models have recently gained attention in speech enhancement (SE), providing an alternative to conventional supervised methods. These models transform clean speech training samples into Gaussian noise, usually centered on noisy speech, and subsequently learn a parameterized…

Cited by 12SourceScholar
2024

Posterior Sampling Algorithms for Unsupervised Speech Enhancement with Recurrent Variational Autoencoder

ICASSP 2024accepted

In this paper, we address the unsupervised speech enhancement problem based on recurrent variational autoencoder (RVAE). This approach offers promising generalization performance over the supervised counterpart. Nevertheless, the involved iterative variational expectation-maximization (VEM) process…

Cited by 0SourceScholar
2024

Unsupervised Speech Enhancement with Diffusion-Based Generative Models

ICASSP 2024accepted

Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when generalising to unseen conditions. To address this issue, we introduce an alterna…

Cited by 0SourceScholar
2023

Audio-Visual Speech Enhancement with a Deep Kalman Filter Generative Model

ICASSP 2023accepted

Deep latent variable generative models based on variational autoencoder (VAE) have shown promising performance for audio-visual speech enhancement (AVSE). The underlying idea is to learn a VAE-based audio-visual prior distribution for clean speech data, and then combine it with a statistical noise m…

Cited by 0SourceScholar
2022

The Impact of Removing Head Movements on Audio-Visual Speech Enhancement

ICASSP 2022accepted

This paper investigates the impact of head movements on audio-visual speech enhancement (AVSE). Although being a common conversational feature, head movements have been ignored by past and recent studies: they challenge today’s learning-based methods as they often degrade the performance of models t…

Cited by 0SourceScholar
2021

Switching Variational Auto-Encoders for Noise-Agnostic Audio-Visual Speech Enhancement

ICASSP 2021accepted

Recently, audio-visual speech enhancement has been tackled in the unsupervised settings based on variational auto-encoders (VAEs), where during training only clean data is used to train a generative model for speech, which at test time is combined with a noise model, e.g. nonnegative matrix factoriz…

Cited by 0SourceScholar
2020

Low Mutual and Average Coherence Dictionary Learning Using Convex Approximation

ICASSP 2020accepted

In dictionary learning, a desirable property for the dictionary is to be of low mutual and average coherences. Mutual coherence is defined as the maximum absolute correlation between distinct atoms of the dictionary, whereas the average coherence is a measure of the average correlations. In this pap…

Cited by 3SourceScholar
2020

Robust Unsupervised Audio-Visual Speech Enhancement Using a Mixture of Variational Autoencoders

ICASSP 2020accepted

Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised speech enhancement. When visual data is clean, speech enhancement with audio-vi…

Cited by 0SourceScholar