DSPAST: DISENTANGLED REPRESENTATIONS FOR SPATIAL AUDIO REASONING WITH LARGE LANGUAGE MODELS
Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the type of sound events, as well as the direction and distance of…