Quality and Complexity Tradeoffs for DNN-Based Binaural Speech Enhancement in Hearing Aids
Parth Mishra, Deepak Kadetotad, Eric A. Durant, Terence Betlehem, Martin F. McKinney
Abstract
This paper investigates the impact of input feature frequency resolution and hyperparameter choices on deep neural network based binaural speech enhancement (SE) for hearing aids. We analyze how time-frequency resolution parameters, such as window length and hop ratio, influence speech intelligibility and computational efficiency. Performance metrics, including Modified Binaural Short-Time Objective Intelligibility (MBSTOI) and Signal-to-Noise Ratio improvement (SNRi) are used to evaluate model efficacy. Our findings show that higher frequency resolutions improve intelligibility but demand greater computational resources, while optimized window and hop settings strike a balance between performance and complexity. Additionally, we use Pareto curves to analyze trade-offs between model complexity and performance, offering guidelines for designing efficient audio pipelines for resource-constrained devices, paving the way for practical binaural SE deployment in hearing aids.
BibTeX
@inproceedings{icassp2025_qualityandcomple,
title = {Quality and Complexity Tradeoffs for DNN-Based Binaural Speech Enhancement in Hearing Aids},
author = {Parth Mishra and Deepak Kadetotad and Eric A. Durant and Terence Betlehem and Martin F. McKinney},
booktitle = {ICASSP 2025},
year = {2025}
}