← Search

Veronica Chelu

3 accepted papers

2022

A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature Predictions

AAAI 2022technical

Estimating value functions is a core component of reinforcement learning algorithms. Temporal difference (TD) learning algorithms use bootstrapping, i.e. they update the value function toward a learning target using value estimates at subsequent time-steps. Alternatively, the value function can be u…

Cited by 1SourcePDFScholar
2022

Learning Expected Emphatic Traces for Deep RL

AAAI 2022technical

Off-policy sampling and experience replay are key for improving sample efficiency and scaling model-free temporal difference learning methods. When combined with function approximation, such as neural networks, this combination is known as the deadly triad and is potentially unstable. Recently, it h…

Cited by 16SourcePDFScholar