2025
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
AISTATS 2025poster
This paper proposes a computationally tractable algorithm for learning infinite-horizon average-reward linear mixture Markov decision processes (MDPs) under the Bellman optimality condition. Our algorithm for linear mixture MDPs achieves a nearly minimax optimal regret upper bound of $\widetilde{\ma…