2024
Margin Matching Preference Optimization: Enhanced Model Alignment with Granular Feedback
EMNLP 2024finding
Large language models (LLMs) fine-tuned with alignment techniques, such as reinforcement learning from human feedback, have been instrumental in developing some of the most capable AI systems to date. Despite their success, existing methods typically rely on simple binary labels, such as those indic…