← Search

Keith Ross

3 accepted papers

2020

BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning

NeurIPS 2020poster

There has recently been a surge in research in batch Deep Reinforcement Learning (DRL), which aims for learning a high-performing policy from a given dataset without additional interactions with the environment. We propose a new algorithm, Best-Action Imitation Learning (BAIL), which strives for bot…

2020

Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling

ICML 2020poster

We aim to develop off-policy DRL algorithms that not only exceed state-of-the-art performance but are also simple and minimalistic. For standard continuous control benchmarks, Soft Actor-Critic (SAC), which employs entropy maximization, currently provides state-of-the-art performance. We first demon…