← Search

Haoxing Tian

3 accepted papers

2026

Bridging the Gap Between Average and Discounted TD Learning

ICML 2026poster

The analysis of Temporal Difference (TD) learning in the average-reward setting faces notable theoretical difficulties because the Bellman operator is not contractive with respect to any norm. This complicates standard analyses of stochastic updates that are effective in discounted settings. Althoug…

Cited by 0SourceScholar
2023

Convergence of Actor-Critic with Multi-Layer Neural Networks

NeurIPS 2023poster

The early theory of actor-critic methods considered convergence using linear function approximators for the policy and value functions. Recent work has established convergence using neural network approximators with a single hidden layer. In this work we are taking the natural next step and establis…

Cited by 5SourcePDFScholar
2023

On the Performance of Temporal Difference Learning With Neural Networks

ICLR 2023poster

Neural Temporal Difference (TD) Learning is an approximate temporal difference method for policy evaluation that uses a neural network for function approximation. Analysis of Neural TD Learning has proven to be challenging. In this paper we provide a convergence analysis of Neural TD Learning with a…

Cited by 9SourcePDFScholar