2026
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
ICLR 2026poster
Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model. However, much less is known about the constrained average-reward MDP (CAMDP), where policies must satisfy long-run averag…