← Search

Rahul Singh

7 accepted papers

2026

Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning

AAAI 2026technical

We study infinite-horizon average-reward reinforcement learning for continuous space Lipschitz Markov decision processes (MDPs) in which an agent can play policies from a given set Φ. The proposed algorithms efficiently explore the policy space by “zooming” into the “promising regions” of Φ, thereby

Cited by 0SourcePDFScholar
2025

Latent Representation Learning for Multimodal Brain Activity Translation

ICASSP 2025accepted

Neuroscience employs diverse neuroimaging techniques, each offering distinct insights into brain activity, from electrophysiological recordings such as EEG, which have high temporal resolution, to hemodynamic modalities such as fMRI, which have increased spatial precision. However, integrating these…

Cited by 0SourceScholar
2024

Finite Time Logarithmic Regret Bounds for Self-Tuning Regulation

ICML 2024poster

We establish the first finite-time logarithmic regret bounds for the self-tuning regulation problem. We introduce a modified version of the certainty equivalence algorithm, which we call PIECE, that clips inputs in addition to utilizing probing inputs for exploration. We show that it has a $C \log T…

Cited by 0SourcePDFScholar
2022

Augmented RBMLE-UCB Approach for Adaptive Control of Linear Quadratic Systems

NeurIPS 2022accept

We consider the problem of controlling an unknown stochastic linear system with quadratic costs -- called the adaptive LQ control problem. We re-examine an approach called ``Reward-Biased Maximum Likelihood Estimate'' (RBMLE) that was proposed more than forty years ago, and which predates the ``Uppe…

Cited by 13SourcePDFScholar
2022

Reinforcement Learning Augmented Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits

AAAI 2022technical

We study a finite-horizon restless multi-armed bandit problem with multiple actions, dubbed as R(MA)^2B. The state of each arm evolves according to a controlled Markov decision process (MDP), and the reward of pulling an arm depends on both the current state and action of the corresponding MDP. Sinc…

Cited by 23SourcePDFScholar