2025
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits
ICLR 2025poster
Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only historical logged data. However, existing methods in OPE/L cannot handle many challenging but prevalent scenarios such…