2025
PORT: Preference Optimization on Reasoning Traces
NAACL 2025long
Preference optimization methods have been successfully applied to improve not only the alignment of large language models (LLMs) with human values, but also specific natural language tasks such as summarization and stylistic continuations. This paper proposes using preference optimization methods on…