2025
RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation
EMNLP 2025
Large language models (LLMs) possess strong multilingual capabilities, and combining Reinforcement Learning from Human Feedback (RLHF) with translation tasks has shown great potential. However, we observe that this paradigm performs unexpectedly poorly when applied to colloquial subtitle translation