2025
Infinite-Horizon Reinforcement Learning with Multinomial Logit Function Approximation
AISTATS 2025poster
We study model-based reinforcement learning with non-linear function approximation where the transition function of the underlying Markov decision process (MDP) is given by a multinomial logit (MNL) model. We develop a provably efficient discounted value iteration-based algorithm that works for both…