← Search

Jakub Grudzien

1 accepted papers

2022

Mirror Learning: A Unifying Framework of Policy Optimisation

ICML 2022spotlight

Modern deep reinforcement learning (RL) algorithms are motivated by either the general policy improvement (GPI) or trust-region learning (TRL) frameworks. However, algorithms that strictly respect these theoretical frameworks have proven unscalable. Surprisingly, the only known scalable algorithms v…