← Search

Ziwei Guan

3 accepted papers

2022

PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method

ICLR 2022poster

Emphatic temporal difference (ETD) learning (Sutton et al., 2016) is a successful method to conduct the off-policy value function evaluation with function approximation. Although ETD has been shown to converge asymptotically to a desirable value function, it is well-known that ETD often encounters a…

Cited by 9SourcePDFScholar
2021

When Will Generative Adversarial Imitation Learning Algorithms Attain Global Convergence

AISTATS 2021poster

Generative adversarial imitation learning (GAIL) is a popular inverse reinforcement learning approach for jointly optimizing policy and reward from expert trajectories. A primary question about GAIL is whether applying a certain policy gradient algorithm to GAIL attains a global minimizer (i.e., yie…

Cited by 25SourcePDFScholar