Categorical Distributional Reinforcement Learning with Kullback-Leibler Divergence: Convergence and Asymptotics
We study the problem of distributional reinforcement learning using categorical parametrisations and a KL divergence loss. Previous work analyzing categorical distributional RL has done so using a Cramér distance-based loss, simplifying the analysis but creating a theory-practice gap. We introduce a…