2025
Efficient Multi-Policy Evaluation for Reinforcement Learning
AAAI 2025technical
To unbiasedly evaluate multiple target policies, the dominant approach among RL practitioners is to run and evaluate each target policy separately. However, this evaluation method is far from efficient because samples are not shared across policies, and running target policies to evaluate themselves…