Bilinear Exponential Family of MDPs: Frequentist Regret Bound with Tractable Exploration & Planning
We study the problem of episodic reinforcement learning in continuous state-action spaces with unknown rewards and transitions. Specifically, we consider the setting where the rewards and transitions are modeled using parametric bilinear exponential families. We propose an algorithm, that a) uses pe…