← Search

Zuolin Tu

1 accepted papers

2024

Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate

NeurIPS 2024poster

Real-world decision-making tasks are usually partially observable Markov decision processes (POMDPs), where the state is not fully observable. Recent progress has demonstrated that recurrent reinforcement learning (RL), which consists of a context encoder based on recurrent neural networks (RNNs) fo…