2023
Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk Criterion
NeurIPS 2023poster
Distributional reinforcement learning algorithms have attempted to utilize estimated uncertainty for exploration, such as optimism in the face of uncertainty. However, using the estimated variance for optimistic exploration may cause biased data collection and hinder convergence or performance. In t…