← Search

Jacob Buckman

6 accepted papers

2022

When does return-conditioned supervised learning work for offline reinforcement learning?

NeurIPS 2022accept

Several recent works have proposed a class of algorithms for the offline reinforcement learning (RL) problem that we will refer to as return-conditioned supervised learning (RCSL). RCSL algorithms learn the distribution of actions conditioned on both the state and the return of the trajectory. Then…

2021

The Importance of Pessimism in Fixed-Dataset Policy Optimization

ICLR 2021poster

We study worst-case guarantees on the expected return of fixed-dataset policy optimization algorithms. Our core contribution is a unified conceptual and mathematical framework for the study of algorithms in this regime. This analysis reveals that for naive approaches, the possibility of erroneous va…

2019

DeepMDP: Learning Continuous Latent Space Models for Representation Learning

ICML 2019oral

Many reinforcement learning (RL) tasks provide the agent with high-dimensional observations that can be simplified into low-dimensional continuous states. To formalize this process, we introduce the concept of a \texit{DeepMDP}, a parameterized latent space model that is trained via the minimization…

Cited by 378SourcePDFScholar
2018

Is Generator Conditioning Causally Related to GAN Performance?

ICML 2018oral

Recent work suggests that controlling the entire distribution of Jacobian singular values is an important design consideration in deep learning. Motivated by this, we study the distribution of singular values of the Jacobian of the generator in Generative Adversarial Networks. We find that this Jaco…

Cited by 149SourcePDFScholar
2018

Sample-Efficient Reinforcement Learning with Stochastic Ensemble Value Expansion

NeurIPS 2018oral

There is growing interest in combining model-free and model-based approaches in reinforcement learning with the goal of achieving the high performance of model-free algorithms with low sample complexity. This is difficult because an imperfect dynamics model can degrade the performance of the learnin…

2018

Thermometer Encoding: One Hot Way To Resist Adversarial Examples

ICLR 2018poster

It is well known that it is possible to construct "adversarial examples" for neural networks: inputs which are misclassified by the network yet indistinguishable from true data. We propose a simple modification to standard neural network architectures, thermometer encoding, which significantly incre…

Cited by 771SourcePDFScholar