← Search

Dhruv Malik

6 accepted papers

2023

Weighted Tallying Bandits: Overcoming Intractability via Repeated Exposure Optimality

ICML 2023poster

In human-interactive applications of online learning, a human's preferences or abilities are often a function of the algorithm's recent actions. Motivated by this, a significant line of work has formalized settings where an action's loss is a function of the number of times it was played in the prio…

Cited by 2SourcePDFScholar
2021

Sample Efficient Reinforcement Learning In Continuous State Spaces: A Perspective Beyond Linearity

ICML 2021spotlight

Reinforcement learning (RL) is empirically successful in complex nonlinear Markov decision processes (MDPs) with continuous state spaces. By contrast, the majority of theoretical RL literature requires the MDP to satisfy some form of linear structure, in order to guarantee sample efficient RL. Such…

Cited by 11SourcePDFScholar
2019

Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems

AISTATS 2019poster

We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of a canonical stochastic, two-point, derivative-free method for linear-quadratic systems in which the initial state of the system is drawn at random. In partic…

Cited by 243SourcePDFScholar
2018

An Efficient, Generalized Bellman Update For Cooperative Inverse Reinforcement Learning

ICML 2018oral

Our goal is for AI systems to correctly identify and act according to their human user’s objectives. Cooperative Inverse Reinforcement Learning (CIRL) formalizes this value alignment problem as a two-player game between a human and robot, in which only the human knows the parameters of the reward fu…

Cited by 45SourcePDFScholar