2025
POTEC: Off-Policy Contextual Bandits for Large Action Spaces via Policy Decomposition
ICLR 2025spotlight
We study off-policy learning (OPL) of contextual bandit policies in large discrete action spaces where existing methods -- most of which rely crucially on reward-regression models or importance-weighted policy gradients -- fail due to excessive bias or variance. To overcome these issues in OPL, we p…