Learning-Based Utility Estimation with Application to Speech Enhancement of a Moving Speaker
Jie Zhang, Chengqian Jiang, Yichi Wang, Haoyin Yan, Miao Sun
Abstract
Wireless acoustic sensor network (WASN) has become a useful platform for monitoring acoustic scenes and sound acquisition. It is likely that many acoustic devices have a marginal impact on performance, which facilitates a necessity of optimizing the tradeoff between performance and computational load by microphone subset selection (MSS). It was shown that microphone utility can measure the node-specific importance on sound acquisition, which can potentially guide downstream speech processing. However, existing utility estimation methods were primarily designed in cases with a single static speaker. In this paper, we propose a learning-based utility estimation model for moving sources, allowing for the time-varying informative MSS by simply choosing microphones with larger utilities. The proposed model takes the speech representation extracted by the pre-trained wav2vec2.0 and the DFT magnitudes as raw features and employs an encoder for feature compression. The microphone utility is estimated by a remote decoder using the received feature vectors. Experimental results show that the proposed method can obtain a more accurate utility estimate in terms of Pearson correlation coefficient (PCC) and ordering the estimated utilities turns out a more informative microphone subset for the enhancement of a moving speaker.
BibTeX
@inproceedings{icassp2025_learningbaseduti,
title = {Learning-Based Utility Estimation with Application to Speech Enhancement of a Moving Speaker},
author = {Jie Zhang and Chengqian Jiang and Yichi Wang and Haoyin Yan and Miao Sun},
booktitle = {ICASSP 2025},
year = {2025}
}