ICASSP 2024accepted0 citations

Visually Guided Binaural Audio Generation with Cross-Modal Consistency

Miao Liu, Jing Wang, Xinyuan Qian, Xiang Xie

Abstract

Binaural audio delivers an immersive spatial auditory experience to human listeners, but most existing videos lack binaural audio due to the expertise required for recording environments. Recent studies have been dedicated to converting monaural audio into binaural ones conditioned on the visual inputs. In this paper, we propose a novel audio-visual spatialization network with two added audio decoders, which rely on carefully designed visual features to generate audio outputs for the left and right channels, respectively. In addition, we propose an audio-visual matching loss to further explore the correlation between binaural audio and the scene visual input. Experiment results show that the proposed method outperforms several state-of-the-art binaural audio generation methods on two benchmark datasets FAIR-Play and MUSIC-Stereo. Qualitative results are also presented to demonstrate the effectiveness of the proposed method.

BibTeX
@inproceedings{icassp2024_visuallyguidedbi,
  title = {Visually Guided Binaural Audio Generation with Cross-Modal Consistency},
  author = {Miao Liu and Jing Wang and Xinyuan Qian and Xiang Xie},
  booktitle = {ICASSP 2024},
  year = {2024}
}