Binaural sound generation corresponding to omnidirectional video view using angular region-wise source enhancement
Kenta Niwa, Yuma Koizumi, Kazunori Kobayashi, Hisashi Uematsu
Abstract
Web applications for watching omnidirectional video through head-mounted displays (HMDs) or smartphones have been widely distributed. The goal of this study was to generate binaural sounds corresponding to the user viewpoint. Assuming that a microphone array is used for sound recording, the enhanced signal for each angular region can be extracted. By convolving head-related transfer functions (HRTFs) and enhanced signals and re-synthesizing them, binaural sounds corresponding to the user viewpoint can be virtually generated. In this paper, we propose a method for achieving angular region-wise source enhancement by generating a multichannel Wiener filter based on the power spectral density (PSD)-estimation-in-beamspace method. To measure user localization when watching omnidirectional video through an HMD, we used a system that enables the generation of binaural sounds corresponding to the user viewpoint in real time. Through subjective tests, we confirmed that sound localization corresponding to the user viewpoint can be obtained when applying about a 40-degree angular region-wise source enhancement.
BibTeX
@inproceedings{icassp2016_binauralsoundgen,
title = {Binaural sound generation corresponding to omnidirectional video view using angular region-wise source enhancement},
author = {Kenta Niwa and Yuma Koizumi and Kazunori Kobayashi and Hisashi Uematsu},
booktitle = {ICASSP 2016},
year = {2016}
}