Spatial Audio Coding without Recourse to Background Signal Compression
Abstract
The MPEG-H 3D audio standard applies singular value decomposition (SVD) to the input higher order ambisonics data, then encodes each predominant (foreground) sound component independently using a standard core audio codec. The residual (background) signal is encoded in the ambisonic domain. This paper is motivated by the observations: i) separate coding of SVD components ignores spatial inter channel masking effects; ii) compression in both SVD and ambisonic domains is difficult to perceptually optimize; iii) Only few predominant components are encoded due to the prohibitive side information cost of specifying SVD basis vectors. The proposed coding architecture overcomes the first two concerns by performing all compression in the SVD domain with a masking threshold that is calculated jointly for all encoded components, thereby accounting for cross-component masking. The third shortcoming is circumvented by a novel method for extending a given set of SVD basis vectors at no side information cost, by computing (at both encoder and decoder) basis vectors to span the null space of the transmitted basis vectors. Experimental results provide evidence for substantial objective and subjective gains.
BibTeX
@inproceedings{icassp2019_spatialaudiocodi,
title = {Spatial Audio Coding without Recourse to Background Signal Compression},
author = {Sina Zamani and Kenneth Rose},
booktitle = {ICASSP 2019},
year = {2019}
}