Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
We study how to fine-tune LLMs using user-edit deployment data consisting of a set of context, an agent's response, and user edits. This deployment data is naturally generated by users in applications such as LLMs-based writing assistants and coding agents. The _natural_ origin of user edits makes i…