← Search

Mridul Agarwal

7 accepted papers

2023

On the Global Convergence of Fitted Q-Iteration with Two-layer Neural Network Parametrization

ICML 2023poster

Deep Q-learning based algorithms have been applied successfully in many decision making problems, while their theoretical foundations are not as well understood. In this paper, we study a Fitted Q-Iteration with two-layer ReLU neural network parameterization, and find the sample complexity guarantee…

Cited by 3SourcePDFScholar
2022

Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Primal-Dual Approach

AAAI 2022technical

Reinforcement learning is widely used in applications where one needs to perform sequential decisions while interacting with the environment. The problem becomes more challenging when the decision requirement includes satisfying some safety constraints. The problem is mathematically formulated as co…

Cited by 75SourcePDFScholar
2022

An explore-then-commit algorithm for submodular maximization under full-bandit feedback

UAI 2022poster

We investigate the problem of combinatorial multi-armed bandits with stochastic submodular (in expectation) rewards and full-bandit feedback, where no extra information other than the reward of selected action at each time step $t$ is observed. We propose a simple algorithm, Explore-Then-Commit Gree…

Cited by 24SourcePDFScholar
2022

Regret guarantees for model-based reinforcement learning with long-term average constraints

UAI 2022poster

We consider the problem of constrained Markov Decision Process (CMDP) where an agent interacts with an ergodic Markov Decision Process. At every interaction, the agent obtains a reward and incurs $K$ costs. The agent aims to maximize the long-term average reward while simultaneously keeping the $K$…

Cited by 20SourcePDFScholar
2021

DART: Adaptive Accept Reject Algorithm for Non-Linear Combinatorial Bandits

AAAI 2021technical

We consider the bandit problem of selecting K out of N arms at each time step. The joint reward can be a non-linear function of the rewards of the selected individual arms. The direct use of a multi-armed bandit algorithm requires choosing among all possible combinations, making the action space lar…

Cited by 11SourcePDFScholar
2021

DESERTS: DElay-tolerant SEmi-autonomous Robot Teleoperation for Surgery

ICRA 2021poster

Telesurgery can be hindered by high-latency and low-bandwidth communication networks, often found in austere settings. Even delays of less than one second are known to negatively impact surgeries. To tackle the effects of connectivity associated with telerobotic surgeries, we propose the DESERTS fra…

Cited by 31SourceScholar