DeepPreNet: A Deep Learning Pre-Processing Method for Speech Distortion Correction in Parametric Array Loudspeaker
Wenyao Ma, Yunxi Zhu, Fengyuan Hao, Liwen Qin, Fengyi Fan, Jun Yang
Abstract
The parametric array loudspeaker produces highly directional sound via a nonlinear process in air, which also introduces inherent baseband distortions. However, conventional recursive modulators employed to compensate for nonlinearity demand substantially increased bandwidth and are not optimized for speech applications. In this paper, we propose a deep learning method tailored for speech, called DeepPreNet. It contains two parts: a pre-processing network (PreNet) and a forward inference model (ForwModel). The ForwModel is a pre-trained network using real recorded speeches to model the actual nonlinear process, enhancing its reliability for PreNet training. The PreNet is trained to generate pre-processed signals, which are subsequently fed into the ForwModel to recover the distortion-free speech. By leveraging the harmonic-rich feature of speech, the proposed method incorporates distortions to reconstruct clean speech, thereby alleviating the bandwidth constraints imposed by the transducer. Experiments in both near- and far-field conditions demonstrate that the proposed method achieves remarkable performance compared to refined baseline techniques with the real transducer response.
BibTeX
@inproceedings{icassp2025_deepprenetadeepl,
title = {DeepPreNet: A Deep Learning Pre-Processing Method for Speech Distortion Correction in Parametric Array Loudspeaker},
author = {Wenyao Ma and Yunxi Zhu and Fengyuan Hao and Liwen Qin and Fengyi Fan and Jun Yang},
booktitle = {ICASSP 2025},
year = {2025}
}