2025
Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol
NeurIPS 2025poster
Holdout validation and hyperparameter tuning from data is a long-standing problem in offline reinforcement learning (RL). A standard framework is to use off-policy evaluation (OPE) methods to evaluate and select the policies, but OPE either incurs exponential variance (e.g., importance sampling) or…