2024
Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate
NeurIPS 2024poster
Real-world decision-making tasks are usually partially observable Markov decision processes (POMDPs), where the state is not fully observable. Recent progress has demonstrated that recurrent reinforcement learning (RL), which consists of a context encoder based on recurrent neural networks (RNNs) fo…