← Search

Ali Aroudi

7 accepted papers

2025

Advancing Active Speaker Detection for Egocentric Videos

ICASSP 2025accepted

This paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing th…

Cited by 0SourceScholar
2025

Reexamining the Efficacy of MetricGAN for Speech Enhancement

ICASSP 2025accepted

MetricGAN, a notable generative approach, provides an effective framework to train speech enhancement models to produce high metric scores. However, we identify two key limitations of current MetricGAN-family models, i.e. neglecting certain mainstream metrics during evaluation and conducting evaluat…

Cited by 0SourceScholar
2021

DBnet: Doa-Driven Beamforming Network for end-to-end Reverberant Sound Source Separation

ICASSP 2021accepted

Many deep learning techniques are available to perform source separation and reduce background noise. However, designing an end-to-end multi-channel source separation method using deep learning and conventional acoustic signal processing techniques still remains challenging. In this paper we propose…

Cited by 0SourceScholar
2020

Improving Auditory Attention Decoding Performance of Linear and Non-Linear Methods using State-Space Model

ICASSP 2020accepted

Identifying the target speaker in hearing aid applications is crucial to improve speech understanding. Recent advances in electroencephalography (EEG) have shown that it is possible to identify the target speaker from single-trial EEG recordings using auditory attention decoding (AAD) methods. AAD m…

Cited by 0SourceScholar
2018

EEG-Based Auditory Attention Decoding Using Steerable Binaural Superdirective Beamformer

ICASSP 2018accepted

During the last decades significant progress in multi-microphone speech enhancement algorithms has been made for hearing aids. However, the performance of many algorithms depends on identifying the target speaker to be enhanced. To identify the target speaker from single-trial EEG recordings in an a…

Cited by 0SourceScholar
2016

Auditory attention decoding with EEG recordings using noisy acoustic reference signals

ICASSP 2016accepted

To decode auditory attention from electroencephalography (EEG) recordings in a cocktail-party scenario with two competing speakers a least-squares method has recently been proposed, showing a promising decoding accuracy. This method however requires the clean speech signals of both the attended and…

Cited by 0SourceScholar