2023
Aligning Language Models with Preferences through $f$-divergence Minimization
ICML 2023poster
Aligning language models with preferences can be posed as approximating a target distribution representing some desired behavior. Existing approaches differ both in the functional form of the target distribution and the algorithm used to approximate it. For instance, Reinforcement Learning from Huma…