2023
Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive Approach
NeurIPS 2023poster
Policy evaluation via Monte Carlo (MC) simulation is at the core of many MC Reinforcement Learning (RL) algorithms (e.g., policy gradient methods). In this context, the designer of the learning system specifies an interaction budget that the agent usually spends by collecting trajectories of *fixed…