← Search

Ah Jeong Seo

1 accepted papers

2024

Margin Matching Preference Optimization: Enhanced Model Alignment with Granular Feedback

EMNLP 2024finding

Large language models (LLMs) fine-tuned with alignment techniques, such as reinforcement learning from human feedback, have been instrumental in developing some of the most capable AI systems to date. Despite their success, existing methods typically rely on simple binary labels, such as those indic…