A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
We study reinforcement learning in infinite-horizon average-reward settings with linear MDPs. Previous work addresses this problem by approximating the average-reward setting by discounted setting and employing a value iteration-based algorithm that uses clipping to constrain the span of the value f…