← Search

Anupam Nayak

1 accepted papers

2026

Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov Games

ICML 2026poster

Reverse Kullback–Leibler (KL) divergence-based regularization with respect to a fixed reference policy is widely used in modern reinforcement learning to preserve the desired traits of the reference policy and sometimes to promote exploration (using uniform reference policy, known as entropy regular…

Cited by 0SourceScholar