← Search

Avik Kar

3 accepted papers

2026

Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning

AAAI 2026technical

We study infinite-horizon average-reward reinforcement learning for continuous space Lipschitz Markov decision processes (MDPs) in which an agent can play policies from a given set Φ. The proposed algorithms efficiently explore the policy space by “zooming” into the “promising regions” of Φ, thereby

Cited by 0SourcePDFScholar
2024

Finite Time Logarithmic Regret Bounds for Self-Tuning Regulation

ICML 2024poster

We establish the first finite-time logarithmic regret bounds for the self-tuning regulation problem. We introduce a modified version of the certainty equivalence algorithm, which we call PIECE, that clips inputs in addition to utilizing probing inputs for exploration. We show that it has a $C \log T…

Cited by 0SourcePDFScholar