2026
PAC-Bayesian Reinforcement Learning Trains Generalizable Policies
ICML 2026poster
We derive a novel PAC-Bayesian generalization bound for reinforcement learning that explicitly accounts for Markov dependencies in the data, through the chain's mixing time. This contributes to overcoming challenges in obtaining generalization guarantees for reinforcement learning, where the sequent…