2026
Efficient Offline Reinforcement Learning via Peer-Influenced Constraint
ICLR 2026poster
Offline reinforcement learning (RL) seeks to learn an optimal policy from a fixed dataset, but distributional shift between the dataset and the learned policy often leads to suboptimal real-world performance. Existing methods typically use behavior policy regularization to constrain the learned poli…