2025
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
NeurIPS 2025poster
We study offline reinforcement learning in average-reward MDPs, which presents increased challenges from the perspectives of distribution shift and non-uniform coverage, and has been relatively underexamined from a theoretical perspective. While previous work obtains performance guarantees under sin…