← Search

Brahma Pavse

1 accepted papers

2020

Reducing Sampling Error in Batch Temporal Difference Learning

ICML 2020poster

Temporal difference (TD) learning is one of the main foundations of modern reinforcement learning. This paper studies the use of TD(0), a canonical TD algorithm, to estimate the value function of a given policy from a batch of data. In this batch setting, we show that TD(0) may converge to an inaccu…

Cited by 16SourcePDFScholar