ICLR 2022poster21 citations

Topological Experience Replay

Zhang-Wei Hong, Tao Chen, Yen-Chen Lin, Joni Pajarinen, Pulkit Agrawal

Abstract

State-of-the-art deep Q-learning methods update Q-values using state transition tuples sampled from the experience replay buffer. This strategy often randomly samples or prioritizes data sampling based on measures such as the temporal difference (TD) error. Such sampling strategies can be inefficient at learning Q-function since a state's correct Q-value preconditions on the accurate successor states' Q-value. Disregarding such a successor's value dependency leads to useless updates and even learning wrong values. To expedite Q-learning, we maintain states' dependency by organizing the agent's experience into a graph. Each edge in the graph represents a transition between two connected states. We perform value backups via a breadth-first search that expands vertices in the graph starting from the set of terminal states successively moving backward. We empirically show that our method is substantially more data-efficient than several baselines on a diverse range of goal-reaching tasks. Notably, the proposed method also outperforms baselines that consume more batches of training experience.

Deep reinforcement learningexperience replay
BibTeX
@inproceedings{
hong2022topological,
title={Topological Experience Replay},
author={Zhang-Wei Hong and Tao Chen and Yen-Chen Lin and Joni Pajarinen and Pulkit Agrawal},
booktitle={International Conference on Learning Representations},
year={2022},
url={https://openreview.net/forum?id=OXRZeMmOI7a}
}
Topological Experience Replay · ICLR 2022