← Search

Michael Gillhofer

1 accepted papers

2019

RUDDER: Return Decomposition for Delayed Rewards

NeurIPS 2019poster

We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate reward plus the expected future rewards. The latter are related to bias problems in temporal difference (TD) learning an…