2026
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
ICASSP 2026poster
While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one bottleneck is the lack of large-scale chain-of-thought audio…