2025
Revisiting a Design Choice in Gradient Temporal Difference Learning
ICLR 2025poster
Off-policy learning enables a reinforcement learning (RL) agent to reason counterfactually about policies that are not executed and is one of the most important ideas in RL. It, however, can lead to instability when combined with function approximation and bootstrapping, two arguably indispensable i…