← Search

Blake A. Richards

1 accepted papers

2022

A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature Predictions

AAAI 2022technical

Estimating value functions is a core component of reinforcement learning algorithms. Temporal difference (TD) learning algorithms use bootstrapping, i.e. they update the value function toward a learning target using value estimates at subsequent time-steps. Alternatively, the value function can be u…

Cited by 1SourcePDFScholar