ICASSP 2026oral0 citations

AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs

Townim Faisal Chowdhury, Siqi Pan, Zhibin Liao

Abstract

Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability is that individual neurons in these models frequently activate in response to several unrelated concepts. We introduce the first mechanistic interpretability framework for AudioLLMs, leveraging sparse autoencoders (SAEs) to disentangle polysemantic activations into monosemantic features. Our pipeline identifies representative audio clips, assigns meaningful names via automated captioning, and validates concepts through human evaluation and steering. Experiments show that AudioLLMs encode structured and interpretable features, enhancing transparency and control. This work provides a foundation for trustworthy deployment in high-stakes domains and enables future extensions to larger models, multilingual audio, and more fine-grained paralinguistic features. Project URL: https://townim-faisal.github.io/AutoInterpret-AudioLLM/

BibTeX
@inproceedings{icassp2026_ardaframeworkfor,
  title = {AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs},
  author = {Townim Faisal Chowdhury and Siqi Pan and Zhibin Liao},
  booktitle = {ICASSP 2026},
  year = {2026}
}
AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs · ICASSP 2026