2020
A Local Temporal Difference Code for Distributional Reinforcement Learning
NeurIPS 2020poster
Recent theoretical and experimental results suggest that the dopamine system implements distributional temporal difference backups, allowing learning of the entire distributions of the long-run values of states rather than just their expected values. However, the distributional codes explored so far…