A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic Approach
This work examines average-reward reinforcement learning with general policy parametrization. Existing state-of-the-art (SOTA) guarantees for this problem are either suboptimal or hindered by several challenges, including poor scalability with respect to the size of the state-action space, high iter…