← Search

Nilesh Madhu

6 accepted papers

2024

Phase Reconstruction in Single Channel Speech Enhancement Based on Phase Gradients and Estimated Clean-Speech Amplitudes

ICASSP 2024accepted

Phase gradients can help enforce phase consistency across time and frequency, further improving the output of speech enhancement approaches. Recently, neural networks were used to estimate the phase gradients from the short-term amplitude spectra of clean speech. These were then used to synthesise p…

Cited by 0SourceScholar
2023

Aiding Speech Harmonic Recovery in DNN-Based Single Channel Noise Reduction Using Cepstral Excitation Manipulation (CEM) Components

ICASSP 2023accepted

Weak harmonics of voiced speech segments are often lost during the process of noise suppression – especially at low SNRs. This leads to a distortion in the harmonic structure, and an accompanying loss in quality. In this paper, inspired by previous work on speech harmonic enhancement using statistic…

Cited by 0SourceScholar
2023

Exploiting Speaker Embeddings for Improved Microphone Clustering and Speech Separation in ad-hoc Microphone Arrays

ICASSP 2023accepted

For separating sources captured by ad hoc distributed microphones a key first step is assigning the microphones to the appropriate source-dominated clusters. The features used for such (blind) clustering are based on a fixed length embedding of the audio signals in a high-dimensional latent space. I…

Cited by 0SourceScholar
2023

Improved Deep Speaker Localization and Tracking: Revised Training Paradigm and Controlled Latency

ICASSP 2023accepted

Even without a separate tracking algorithm, the directions of arrival (DOAs) of moving talkers can be estimated with a deep neural network (DNN) when the movement trajectories used for training allow the generalization to real signals. Previously, we proposed a framework for generating training data…

Cited by 0SourceScholar
2023

Margin-Mixup: A Method for Robust Speaker Verification In Multi-Speaker Audio

ICASSP 2023accepted

This paper is concerned with the task of speaker verification on audio with multiple overlapping speakers. Most speaker verification systems are designed with the assumption of a single speaker being present in a given audio segment. However, in a real-world setting this assumption does not always h…

Cited by 0SourceScholar
2020

Least-Squares DOA Estimation with an Informed Phase Unwrapping and Full Bandwidth Robustness

ICASSP 2020accepted

The weighted least-squares (WLS) direction-of-arrival estimator that minimizes an error based on interchannel phase differences is both computationally simple and flexible. However, the approach has several limitations, including an inability to cope with spatial aliasing and a sensitivity to phase…

Cited by 0SourceScholar