2026
The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMs
ICASSP 2026poster
Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks d…