Ultra low-compute complex spectral masking for multichannel speech enhancement
Ashutosh Pandey, Juan Azcarreta
Abstract
We present a streamlined framework for complex spectral masking that processes multichannel speech with minimal computational demands, enhancing both spectral magnitude and phase by integrating low-compute models with the Multi-Channel Wiener Filter (MCWF). Our methodology employs a two-stage, end-to-end training approach where a deep neural network (DNN) first estimates MCWF weights, followed by another DNN that refines the MCWF output, enhancing spectral masking quality. This architecture not only outperforms the traditional oracle Minimum Variance Distortionless Response (MVDR) beamformer but also maintains high efficiency, requiring less than 50MMACs for processing one second of 8-channel audio. Empirical results demonstrate that our framework exceeds the performance of existing low-compute models, offering significant enhancements with minimal computational demands, making it ideal for deployment on edge devices with limited computational resources.
BibTeX
@inproceedings{icassp2025_ultralowcomputec,
title = {Ultra low-compute complex spectral masking for multichannel speech enhancement},
author = {Ashutosh Pandey and Juan Azcarreta},
booktitle = {ICASSP 2025},
year = {2025}
}