Dual Position Attention Time-Frequency Network for Binaural Audio Synthesis
Changjun He, Weiping Chen, Mingjiang Wang
Abstract
In applications such as virtual reality and augmented reality, binaural audio provides listeners with a more immersive experience. To synthesize binaural audio with enhanced spatial localization, especially in scenarios involving moving sound sources, accurate phase estimation is crucial. However, existing deep learning methods have yet to achieve satisfactory results in this area. To address this issue, this paper introduces a Dual Position Attention Time-Frequency Network (DPATFNet). Specifically, our approach targets the interaural differences and Doppler effects induced by sound source movement, guiding the synthesis process from monaural to binaural audio through strong positional conditions in the time-frequency domain. The network employs a Dual Position Attention Block (DPAB) to effectively focus on sound source movement and improve phase estimation performance. The proposed DPATFNet demonstrates a strong capability for synthesizing accurate binaural audio, with experimental results on the Binaural Speech dataset showing that DPATFNet achieves state-of-the-art performance in phase metrics (Phase-L2: 0.717, IPD-L2: 1.020, Wave-L2: 0.148, Amplitude-L2: 0.037).
BibTeX
@inproceedings{icassp2025_dualpositionatte,
title = {Dual Position Attention Time-Frequency Network for Binaural Audio Synthesis},
author = {Changjun He and Weiping Chen and Mingjiang Wang},
booktitle = {ICASSP 2025},
year = {2025}
}