AAAI 2026technical0 citations
Measuring the Unmeasurable: Unveiling Latent Cognitive Capabilities of LLM
Cui Danxin, Sihang Jiang, Keyi Wang, Zhiyi Duan, Yanghua Xiao, Bi Yude, Jiaqing Liang, Minggui He
Abstract
As large language models (LLMs) are increasingly deployed in high-stakes domains such as education, healthcare, and law, accurately evaluating their nuanced reasoning process becomes essential to ensure their safety, reliability, and trustworthiness. However, most existing benchmarks evaluate LLMs at a coarse granularity. Current benchmarks lack a unified framework and rely on single‐task datasets, overlooking the intermediate steps of complex reasoning. This results in redundant overlap across benchmarks, poor generalization to multifaceted real-world tasks, and underutilizes the rich reasoning traces generated by advanced LLMs.
BibTeX
@inproceedings{aaai2026_measuringtheunme,
title = {Measuring the Unmeasurable: Unveiling Latent Cognitive Capabilities of LLM},
author = {Cui Danxin and Sihang Jiang and Keyi Wang and Zhiyi Duan and Yanghua Xiao and Bi Yude and Jiaqing Liang and Minggui He and Shimin Tao and Yilun Liu},
booktitle = {AAAI 2026},
year = {2026}
}