ICASSP 2025accepted0 citations

Robust Fusion of Bone and Air-Conducted Sensors for Speech Enhancement with Adaptive Temporal-Frequency Attention

Zhenglong Liu, Zhe Chen, Fuliang Yin

Abstract

The multi-modal speech enhancement method has improved performance due to the diverse sources of its input data, which includes low-distortion air-conducted (AC) signals and low-noise bone-conducted (BC) signals. In light of these considerations, a novel complex-domain deep learning-based BC and AC speech fusion enhancement method is proposed. Specifically, the BC and AC spectrum is fed into a pre-fusion module firstly. Then the preliminary fused feature is concatenated with the original signals and fed into encoder, followed by adaptive time-frequency attention modules, in which the self-attention mechanism is employed to reconstruct the encoded features in both the time and frequency axes. After that, two masks are generated by decoders and multiplied with the original BC and AC spectrum. Finally, the two signals are summed and converted to the time-domain. The disabled situation when only one sensor works is considered to improve the robustness by introducing a probability of failure in the process of feeding the data to train the model. Experiments demonstrate that the proposed method outperforms existing fusion approaches, especially in the aspect of robustness against one invalid input channel.

BibTeX
@inproceedings{icassp2025_robustfusionofbo,
  title = {Robust Fusion of Bone and Air-Conducted Sensors for Speech Enhancement with Adaptive Temporal-Frequency Attention},
  author = {Zhenglong Liu and Zhe Chen and Fuliang Yin},
  booktitle = {ICASSP 2025},
  year = {2025}
}