AAAI 2026technical0 citations

Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models

Changyue Wang, Weihang Su, Qingyao Ai, Yiqun Liu

Abstract

Large Reasoning Models (LRMs) extend large language models with explicit, multi-step reasoning traces to enhance transparency and performance on complex tasks. However, these reasoning traces can be redundant or logically inconsistent, becoming a new and hard-to-detect source of hallucination. Existing hallucination detection methods focus primarily on answer-level uncertainty and often fail to detect hallucinations or logical inconsistencies arising from the model’s reasoning trace. This oversight is particularly problematic for LRMs, where the explicit thinking trace is not only an important support to the model

BibTeX
@inproceedings{aaai2026_jointevaluationo,
  title = {Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models},
  author = {Changyue Wang and Weihang Su and Qingyao Ai and Yiqun Liu},
  booktitle = {AAAI 2026},
  year = {2026}
}