2026
Taming the Aleatoric Impulse in Off-Policy Reinforcement Learning
ICML 2026poster
Off-policy reinforcement learning is vulnerable to overestimation bias, which is rooted in the total value uncertainty. However, existing methods typically misaddress this by targeting the epistemic component, neglecting the aleatoric component. We identify for the first time that this oversight fai…