ICLR 2026poster0 citations

AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models

Kai Li, Can Shen, Yile Liu, Jirui Han, Kelong zheng, Xuechao Zou, Lionel Z. WANG, Shun Zhang

Abstract

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) require a rigorous assessment of their trustworthiness. However, existing evaluation frameworks, primarily designed for text, are not equipped to handle the unique vulnerabilities introduced by audio’s acoustic properties. We find that significant trustworthiness risks in ALLMs arise from non-semantic acoustic cues, such as timbre, accent, and background noise, which can be used to manipulate model behavior. To address this gap, we propose AudioTrust, the first framework for large-scale and systematic evaluation of ALLM trustworthiness concerning these audio-specific risks. AudioTrust spans six key dimensions: fairness, hallucination, safety, privacy, robustness, and authenticition. It is implemented through 26 distinct sub-tasks and a curated dataset of over 4,420 audio samples collected from real-world scenarios (e.g., daily conversations, emergency calls, and voice assistant interactions), purposefully constructed to probe the trustworthiness of ALLMs across multiple dimensions. Our comprehensive evaluation includes 18 distinct experimental configurations and employs human-validated automated pipelines to objectively and scalably quantify model outputs. Experimental results reveal the boundaries and limitations of 14 state-of-the-art (SOTA) open-source and closed-source ALLMs when confronted with diverse high-risk audio scenarios, thereby offering critical insights into the secure and trustworthy deployment of future audio models. Our platform and benchmark are publicly available at https://anonymous.4open.science/r/AudioTrust-8715/.

Audio Large Language Model
BibTeX
@inproceedings{
li2026audiotrust,
title={AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models},
author={Kai Li and Can Shen and Yile Liu and Jirui Han and Kelong zheng and Xuechao Zou and Lionel Z. WANG and Shun Zhang and Xingjian Du and Hanjun Luo and Yingbin Jin and Xinxin Xing and Ziyang Ma and Yue Liu and YiFan Zhang and Junfeng Fang and Kun Wang and Yibo Yan and Gelei Deng and Haoyang LI and Yiming Li and Xiaobin Zhuang and Tianlong Chen and Qingsong Wen and Tianwei Zhang and Yang Liu and Haibo Hu and Zhizheng Wu and Xiaolin Hu and Eng Siong Chng and Wenyuan Xu and XiaoFeng Wang and Wei Dong and Xinfeng Li},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=E823AY0taq}
}
AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models · ICLR 2026