← Search

Zhouyang Yu

1 accepted papers

2026

Taming the Aleatoric Impulse in Off-Policy Reinforcement Learning

ICML 2026poster

Off-policy reinforcement learning is vulnerable to overestimation bias, which is rooted in the total value uncertainty. However, existing methods typically misaddress this by targeting the epistemic component, neglecting the aleatoric component. We identify for the first time that this oversight fai…

Cited by 0SourceScholar