AAAI 2026technical0 citations

Fine-Tuned LLMs Know They Don’t Know: A Parameter-Efficient Approach to Recovering Honesty

Zeyu Shi, Ziming Wang, Tianyu Chen, Shiqi Gao, Haoyi Zhou, Qingyun Sun, Jianxin Li

Abstract

The honesty of Large Language Models (LLMs) is increasingly important for safe deployment in high-stakes domains. However, this crucial trait is severely undermined by supervised fine-tuning (SFT), a common technique for model specialization. Existing recovery methods rely on data-intensive global parameter adjustments, implicitly assuming that SFT deeply corrupts the models

BibTeX
@inproceedings{aaai2026_finetunedllmskno,
  title = {Fine-Tuned LLMs Know They Don’t Know: A Parameter-Efficient Approach to Recovering Honesty},
  author = {Zeyu Shi and Ziming Wang and Tianyu Chen and Shiqi Gao and Haoyi Zhou and Qingyun Sun and Jianxin Li},
  booktitle = {AAAI 2026},
  year = {2026}
}