2025
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
NeurIPS 2025poster
Reinforcement learning (RL) is increasingly used to align large language models (LLMs). Off-policy methods offer greater implementation simplicity and data efficiency than on-policy techniques, but often result in suboptimal performance. In this work, we study the intermediate range of algorithms be…