2025
Offline RL via Feature-Occupancy Gradient Ascent
AISTATS 2025poster
We study offline Reinforcement Learning in large infinite-horizon discounted Markov Decision Processes (MDPs) when the reward and transition models are linearly realizable under a known feature map. Starting from the classic linear-program formulation of the optimal control problem in MDPs, we devel…