← Search

Alireza Aghaei

1 accepted papers

2025

Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning

ICML 2025spotlight

Off-policy learning and evaluation leverage logged bandit feedback datasets, which contain context, action, propensity score, and feedback for each data point. These scenarios face significant challenges due to high variance and poor performance with low-quality propensity scores and heavy-tailed re…