← Search

Nishant Sharma

2 accepted papers

2026

Enhancing Trustworthiness of Fine-Tuned LLMs via Regularized Subset Selection

ICLR 2026poster

Supervised fine-tuning (SFT) improves large language model (LLM) perplexity but can also degrade trustworthiness—leading to the generation of untruthful, biased, or unsafe content during user interactions. These issues are often traced back to specific phrases or patterns in the training data. Howev…

Cited by 0SourceScholar