← Search

Andrew J Wagenmaker

4 accepted papers

2022

First-Order Regret in Reinforcement Learning with Linear Function Approximation: A Robust Estimation Approach

ICML 2022oral

Obtaining first-order regret bounds—regret bounds scaling not as the worst-case but with some measure of the performance of the optimal policy on a given instance—is a core question in sequential decision-making. While such bounds exist in many settings, they have proven elusive in reinforcement lea…

Cited by 43SourcePDFScholar
2022

Reward-Free RL is No Harder Than Reward-Aware RL in Linear Markov Decision Processes

ICML 2022spotlight

Reward-free reinforcement learning (RL) considers the setting where the agent does not have access to a reward function during exploration, but must propose a near-optimal policy for an arbitrary reward function revealed only after exploring. In the the tabular setting, it is well known that this is…

Cited by 70SourcePDFScholar