2025
CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation
ACL 2025finding
Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Op…