2024
Randomized algorithms and PAC bounds for inverse reinforcement learning in continuous spaces
NeurIPS 2024poster
This work studies discrete-time discounted Markov decision processes with continuous state and action spaces and addresses the inverse problem of inferring a cost function from observed optimal behavior. We first consider the case in which we have access to the entire expert policy and characterize…