← Search

Huaqing. Xiong

5 accepted papers

2022

Deterministic policy gradient: Convergence analysis

UAI 2022poster

The deterministic policy gradient (DPG) method proposed in Silver et al. [2014] has been demonstrated to exhibit superior performance particularly for applications with multi-dimensional and continuous action spaces. However, it remains unclear whether DPG converges, and if so, how fast it converges…

Cited by 25SourcePDFScholar
2021

Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian Sampling

AAAI 2021technical

Despite the wide applications of Adam in reinforcement learning (RL), the theoretical convergence of Adam-type RL algorithms has not been established. This paper provides the first such convergence analysis for two fundamental RL algorithms of policy gradient (PG) and temporal difference (TD) learni…

Cited by 41SourcePDFScholar
2020

Analysis of Q-learning with Adaptation and Momentum Restart for Gradient Descent

IJCAI 2020poster

Existing convergence analyses of Q-learning mostly focus on the vanilla stochastic gradient descent (SGD) type of updates. Despite the Adaptive Moment Estimation (Adam) has been commonly used for practical Q-learning algorithms, there has not been any convergence guarantee provided for Q-learning wi…

Cited by 0SourcePDFScholar