NeurIPS 2020oral14 citations

Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory

Yufeng Zhang, Qi Cai, Zhuoran Yang, Yongxin Chen, Zhaoran Wang

Abstract

Temporal-difference and Q-learning play a key role in deep reinforcement learning, where they are empowered by expressive nonlinear function approximators such as neural networks. At the core of their empirical successes is the learned feature representation, which embeds rich observations, e.g., images and texts, into the latent space that encodes semantic structures. Meanwhile, the evolution of such a feature representation is crucial to the convergence of temporal-difference and Q-learning.

BibTeX
@inproceedings{NEURIPS2020_e3bc4e7f,
 author = {Zhang, Yufeng and Cai, Qi and Yang, Zhuoran and Chen, Yongxin and Wang, Zhaoran},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
 pages = {19680--19692},
 publisher = {Curran Associates, Inc.},
 title = {Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory},
 url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/e3bc4e7f243ebc05d66a0568a3331966-Paper.pdf},
 volume = {33},
 year = {2020}
}
Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory · NeurIPS 2020