Do Retrieval Augmented Language Models Know When They Don’t Know?
Youchao Zhou, Heyan Huang, Yicheng Liu, Rui Dai, Xinglin Wang, Xingchen Zhang, Shumin Shi, Yang Deng
Abstract
Existing large language models (LLMs) occasionally generate plausible yet factually incorrect responses, known as hallucinations. Two main approaches have been proposed to mitigate hallucinations: retrieval-augmented language models (RALMs) and refusal post-training. However, current research predominantly focuses on their individual effectiveness while overlooking the evaluation of the refusal capability of RALMs. Ideally, if RALMs know when they do not know, they should refuse to answer. In this study, we ask the fundamental question: Do RALMs know when they don’t know? Specifically, we investigate three questions. First, are RALMs well calibrated with respect to different internal and external knowledge states? We examine the influence of various factors. Contrary to expectations, when all retrieved documents are irrelevant, RALMs still tend to refuse questions they could have answered correctly. Next, given the model
BibTeX
@inproceedings{aaai2026_doretrievalaugme,
title = {Do Retrieval Augmented Language Models Know When They Don’t Know?},
author = {Youchao Zhou and Heyan Huang and Yicheng Liu and Rui Dai and Xinglin Wang and Xingchen Zhang and Shumin Shi and Yang Deng},
booktitle = {AAAI 2026},
year = {2026}
}