2021
Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy Estimate
AAAI 2021technical
In a reward-free environment, what is a suitable intrinsic objective for an agent to pursue so that it can learn an optimal task-agnostic exploration policy? In this paper, we argue that the entropy of the state distribution induced by finite-horizon trajectories is a sensible target. Especially, we…