2020
Balancing Learning Speed and Stability in Policy Gradient via Adaptive Exploration
AISTATS 2020poster
In many Reinforcement Learning (RL) applications, the goal is to find an optimal deterministic policy. However, most RL algorithms require the policy to be stochastic in order to avoid instabilities and perform a sufficient amount of exploration. Adjusting the level of stochasticity during the learn…