ICASSP 2025accepted0 citations

Metadata-assisted Pose Correction for Immersive Audio Split Rendering

Rishabh Tyagi, Stefan Bruhn, Jeroen Breebaart

Abstract

Split rendering as recently standardized in 3GPP/IVAS allows for offloading a substantial part of a binaural rendering process from a rendering end-device with limited capabilities, such a wearable, to a more capable entity. This is achieved by computing and transmitting pose correction metadata alongside a pre-rendered binaural audio signal, allowing for head pose adjustments of the pre-rendered binaural signal in response to the headtracking data available at the end-device. In addition to reducing renderer complexity at the end-device, the split renderer also resolves potential large motion-to-sound latencies on the interface between the pre-rendering entity and the end-device. This paper describes practical pre-rendering and coding strategies of the standardized method using new metadata-driven analysis and synthesis methods that allow for such low-latency, low complexity pose correction at the end-device. Evaluation results demonstrate substantial complexity benefits, and a quality level close to the "ideal" alternative of performing full decoding and rendering.

BibTeX
@inproceedings{icassp2025_metadataassisted,
  title = {Metadata-assisted Pose Correction for Immersive Audio Split Rendering},
  author = {Rishabh Tyagi and Stefan Bruhn and Jeroen Breebaart},
  booktitle = {ICASSP 2025},
  year = {2025}
}