2019
Distributional Multivariate Policy Evaluation and Exploration with the Bellman GAN
ICML 2019oral
The recently proposed distributional approach to reinforcement learning (DiRL) is centered on learning the distribution of the reward-to-go, often referred to as the value distribution. In this work, we show that the distributional Bellman equation, which drives DiRL methods, is equivalent to a gene…