2025
A General Framework for Off-Policy Learning with Partially-Observed Reward
ICLR 2025poster
Off-policy learning (OPL) in contextual bandits aims to learn a decision-making policy that maximizes the target rewards by using only historical interaction data collected under previously developed policies. Unfortunately, when rewards are only partially observed, the effectiveness of OPL degrades…