IJCAI 20260 citations

A Survey on Actionable Interpretability in Large Language Models

Jie Cai, Mafizur Rahman, James Enouen, Lijun Qian, Yan Liu

Abstract

Large Language Models (LLMs) have become central to modern AI, with interpretability serving as a critical means of investigating the opaque and highly nonlinear mechanisms encoded within billions of parameters and ensuring trustworthy deployment. However, descriptive interpretability approaches for LLMs remain largely post-hoc, illuminating model behavior without providing the actionable leverage needed to influence or adapt model behavior, thereby limiting their practical utility. Recent work has therefore reframed interpretability as an actionable paradigm, shifting the focus from explanation alone toward methods that connect internal mechanisms to model refinement. This survey reviews LLM interpretability through the lens of actionability, presenting a taxonomy of attributional, concept-based, and mechanistic approaches, along with emerging methods tailored to vision–language models (VLMs). We further examine how interpretability supports downstream objectives such as hallucination mitigation, model editing, fairness, and safety. By positioning interpretability as a pathway to better-guided LLM design and practice, this survey outlines key challenges and future directions toward trustworthy and controllable foundation models.

AI Ethics, Trust, Fairnes: Explainability and interpretabilityAI Ethics, Trust, Fairnes: Fairness and diversityAI Ethics, Trust, Fairnes: Trustworthy AI
BibTeX
@inproceedings{ijcai2026_asurveyonactiona,
  title = {A Survey on Actionable Interpretability in Large Language Models},
  author = {Jie Cai and Mafizur Rahman and James Enouen and Lijun Qian and Yan Liu},
  booktitle = {IJCAI 2026},
  year = {2026}
}
A Survey on Actionable Interpretability in Large Language Models · IJCAI 2026