← Search

Mehrdad Moharrami

2 accepted papers

2023

Performance Bounds for Policy-Based Average Reward Reinforcement Learning Algorithms

NeurIPS 2023poster

Many policy-based reinforcement learning (RL) algorithms can be viewed as instantiations of approximate policy iteration (PI), i.e., where policy improvement and policy evaluation are both performed approximately. In applications where the average reward objective is the meaningful performance metri…

Cited by 4SourcePDFScholar