← Search

Miroslav Štrupl

1 accepted papers

2022

Reward-Weighted Regression Converges to a Global Optimum

AAAI 2022technical

Reward-Weighted Regression (RWR) belongs to a family of widely known iterative Reinforcement Learning algorithms based on the Expectation-Maximization framework. In this family, learning at each iteration consists of sampling a batch of trajectories using the current policy and fitting a new policy…