2023
Prediction and Control in Continual Reinforcement Learning
NeurIPS 2023poster
Temporal difference (TD) learning is often used to update the estimate of the value function which is used by RL agents to extract useful policies. In this paper, we focus on value function estimation in continual reinforcement learning. We propose to decompose the value function into two components…