← Search

Junzi Zhang

4 accepted papers

2022

Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods

AAAI 2022technical

When designing algorithms for finite-time-horizon episodic reinforcement learning problems, a common approach is to introduce a fictitious discount factor and use stationary policies for approximations. Empirically, it has been shown that the fictitious discount factor helps reduce variance, and sta…

Cited by 8SourcePDFScholar
2021

Sample Efficient Reinforcement Learning with REINFORCE

AAAI 2021technical

Policy gradient methods are among the most effective methods for large-scale reinforcement learning, and their empirical success has prompted several works that develop the foundation of their global convergence theory. However, prior works have either required exact gradients or state-action visita…

Cited by 130SourcePDFScholar