← Search

Akarsh Kumar

1 accepted papers

2023

Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks

NeurIPS 2023poster

Most reinforcement learning methods rely heavily on dense, well-normalized environment rewards. DreamerV3 recently introduced a model-based method with a number of tricks that mitigate these limitations, achieving state-of-the-art on a wide range of benchmarks with a single set of hyperparameters. T…

Cited by 5SourcePDFScholar