2023
Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks
NeurIPS 2023poster
Most reinforcement learning methods rely heavily on dense, well-normalized environment rewards. DreamerV3 recently introduced a model-based method with a number of tricks that mitigate these limitations, achieving state-of-the-art on a wide range of benchmarks with a single set of hyperparameters. T…