2025
SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models
EMNLP 2025
Existing alignment methods for preference optimization of large language models (LLMs) aim to enhance model performance by utilizing pairs of positive and negative samples. However, due to the limited capacity of models in scoring or generating responses, the quality of positive and negative samples