NAT3DSound: 3D Spatial Sound Field Synthesis with Multi-Modal Non-Autoregressive Transformer
Fuming You, Rongjie Huang, Boyang Zhang, Yongqi Wang, Zhiqing Hong, Zhimeng Zhang, Zhou Zhao
Abstract
3D spatial sound field synthesis takes the head-mounted audio signals and body poses as input and renders a 3D sound field around the center body, in which spatial audio can be inferred at any arbitrary position. To achieve this, a multi-modal system is required to spatialize input audio signals with the guidance of pose streams. Then, a numerical method is used to render the sound field with hundreds of spatial audio. However, there exist several challenges including hybrid audio input signals, cross-modal matching, and few data resources. To address them, we propose NAT3DSound, a unit-based framework for high-quality 3D spatial sound synthesis, which consists of 1) a general tokenization scheme for both dense and impulsive input audio signals, 2) a non-autoregressive transformer with multi-scale modality fusion for efficient cross-modal alignment and 3) a parallel sampling strategy for fast prediction. Furthermore, we investigate the feasibility of acoustic pre-training for low-resource learning in data-scarce scenarios. Extensive experiments and ablation studies demonstrate the effectiveness of NAT3DSound in terms of spatialization quality and generalization ability. Audio samples are available at http://NAT3DSound.github.io
BibTeX
@inproceedings{icassp2025_nat3dsound3dspat,
title = {NAT3DSound: 3D Spatial Sound Field Synthesis with Multi-Modal Non-Autoregressive Transformer},
author = {Fuming You and Rongjie Huang and Boyang Zhang and Yongqi Wang and Zhiqing Hong and Zhimeng Zhang and Zhou Zhao},
booktitle = {ICASSP 2025},
year = {2025}
}