← Search

Andrey Volozin

1 accepted papers

2025

DPL: Diverse Preference Learning Without A Reference Model

NAACL 2025long

In direct preference alignment in LLMs, most existing methods seek to retrieve the reward function directly from preference data. However, real-world preference data often contains diversity in preference annotations reflective of true human preferences. Existing algorithms, including KTO, do not di…

Cited by 0SourcePDFScholar