2025
Weak-to-Strong Honesty Alignment via Learning-to-Rank Supervision
ACL 2025finding
Honest alignment refers to the ability of a language model to truthfully convey its knowledge limitations by appropriately refusing to answer questions when it lacks sufficient information. Existing solutions, such as prompt engineering and fine-tuning, face limitations: the former provides only mar…