2021
Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative Model
NeurIPS 2021poster
The curse of dimensionality is a widely known issue in reinforcement learning (RL). In the tabular setting where the state space $\mathcal{S}$ and the action space $\mathcal{A}$ are both finite, to obtain a near optimal policy with sampling access to a generative model, the minimax optimal sample co…