2022
Controlling Underestimation Bias in Reinforcement Learning via Quasi-median Operation
AAAI 2022technical
How to get a good value estimation is one of the key problems in reinforcement learning (RL). Current off-policy methods, such as Maxmin Q-learning, TD3 and TADD, suffer from the underestimation problem when solving the overestimation problem. In this paper, we propose the Quasi-Median Operation, a…