2017
Natural Value Approximators: Learning when to Trust Past Estimates
NeurIPS 2017spotlight
Neural networks have a smooth initial inductive bias, such that small changes in input do not lead to large changes in output. However, in reinforcement learning domains with sparse rewards, value functions have non-smooth structure with a characteristic asymmetric discontinuity whenever rewards arr…