2024
Differentially Private No-regret Exploration in Adversarial Markov Decision Processes
UAI 2024poster
We study learning adversarial Markov decision process (MDP) in the episodic setting under the constraint of differential privacy (DP). This is motivated by the widespread applications of reinforcement learning (RL) in non-stationary and even adversarial scenarios, where protecting users’ sensitive i…