AAAI 2026technical0 citations
Fine-Tuned LLMs Know They Don’t Know: A Parameter-Efficient Approach to Recovering Honesty
Zeyu Shi, Ziming Wang, Tianyu Chen, Shiqi Gao, Haoyi Zhou, Qingyun Sun, Jianxin Li
Abstract
The honesty of Large Language Models (LLMs) is increasingly important for safe deployment in high-stakes domains. However, this crucial trait is severely undermined by supervised fine-tuning (SFT), a common technique for model specialization. Existing recovery methods rely on data-intensive global parameter adjustments, implicitly assuming that SFT deeply corrupts the models
BibTeX
@inproceedings{aaai2026_finetunedllmskno,
title = {Fine-Tuned LLMs Know They Don’t Know: A Parameter-Efficient Approach to Recovering Honesty},
author = {Zeyu Shi and Ziming Wang and Tianyu Chen and Shiqi Gao and Haoyi Zhou and Qingyun Sun and Jianxin Li},
booktitle = {AAAI 2026},
year = {2026}
}