2025
Self-Evolutionary Large Language Models Through Uncertainty-Enhanced Preference Optimization
AAAI 2025technical
Iterative preference optimization has recently become one of the de-facto training paradigms for large language models (LLMs), but the performance is still underwhelming due to too much noisy preference data yielded in the loop. To combat this issue, we present an Uncertainty-enhanced Preference Opt…