2022
SURF: Semantic-level Unsupervised Reward Function for Machine Translation
NAACL 2022long
The performance of Reinforcement Learning (RL) for natural language tasks including Machine Translation (MT) is crucially dependent on the reward formulation. This is due to the intrinsic difficulty of the task in the high-dimensional discrete action space as well as the sparseness of the standard r…