2023
Adaptively Calibrated Critic Estimates for Deep Reinforcement Learning
RA-L 2023
Accurate value estimates are important for off-policy reinforcement learning. Algorithms based on temporal difference learning typically are prone to an over- or underestimation bias building up over time. In this letter, we propose a general method called Adaptively Calibrated Critics (ACC) that us