2026
Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov Games
ICML 2026poster
Reverse Kullback–Leibler (KL) divergence-based regularization with respect to a fixed reference policy is widely used in modern reinforcement learning to preserve the desired traits of the reference policy and sometimes to promote exploration (using uniform reference policy, known as entropy regular…