2026
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
ICASSP 2026poster
While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well as the lack of structured intermediate supervision. To address this, we propose a unified knowledge distillation framew…