← Search

Rupak Majumdar

6 accepted papers

2023

Markov Decision Processes with Time-Varying Geometric Discounting

AAAI 2023technical

Canonical models of Markov decision processes (MDPs) usually consider geometric discounting based on a constant discount factor. While this standard modeling approach has led to many elegant results, some recent studies indicate the necessity of modeling time-varying discounting in certain applicati…

Cited by 2SourcePDFScholar
2023

Online Reinforcement Learning with Uncertain Episode Lengths

AAAI 2023technical

Existing episodic reinforcement algorithms assume that the length of an episode is fixed across time and known a priori. In this paper, we consider a general framework of episodic reinforcement learning when the length of each episode is drawn from a distribution. We first establish that this prob…

Cited by 7SourcePDFScholar
2021

Choosing the Initial State for Online Replanning

AAAI 2021technical

The need to replan arises in many applications. However, in the context of planning as heuristic search, it raises an annoying problem: if the previous plan is still executing, what should the new plan search take as its initial state? If it were possible to accurately predict how long replanning wo…

Cited by 3SourcePDFScholar