2026
Enhancing Trustworthiness of Fine-Tuned LLMs via Regularized Subset Selection
ICLR 2026poster
Supervised fine-tuning (SFT) improves large language model (LLM) perplexity but can also degrade trustworthiness—leading to the generation of untruthful, biased, or unsafe content during user interactions. These issues are often traced back to specific phrases or patterns in the training data. Howev…