← Search

Harsh Satija

4 accepted papers

2021

Locally Persistent Exploration in Continuous Control Tasks with Sparse Rewards

ICML 2021spotlight

A major challenge in reinforcement learning is the design of exploration strategies, especially for environments with sparse reward structures and continuous state and action spaces. Intuitively, if the reinforcement signal is very scarce, the agent should rely on some form of short-term memory in o…

2021

Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs

NeurIPS 2021poster

We study the problem of Safe Policy Improvement (SPI) under constraints in the offline Reinforcement Learning (RL) setting. We consider the scenario where: (i) we have a dataset collected under a known baseline policy, (ii) multiple reward signals are received from the environment inducing as many o…

Cited by 24SourcePDFScholar
2019

Randomized Value Functions via Multiplicative Normalizing Flows

UAI 2019poster

Randomized value functions offer a promising approach towards the challenge of efficient exploration in complex environments with high dimensional state and action spaces. Unlike traditional point estimate methods, randomized value functions maintain a posterior distribution over action-space values…

Cited by 47SourcePDFScholar