← Search

Bob McGrew

5 accepted papers

2022

GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

ICML 2022spotlight

Diffusion models have recently been shown to generate high-quality synthetic images, especially when paired with a guidance technique to trade off diversity for fidelity. We explore diffusion models for the problem of text-conditional image synthesis and compare two different guidance strategies: CL…

2020

Emergent Tool Use From Multi-Agent Autocurricula

ICLR 2020spotlight

Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordi…

Cited by 961SourcecodeScholar
2018

Domain Randomization and Generative Models for Robotic Grasping

IROS 2018poster

Deep learning-based robotic grasping has made significant progress thanks to algorithmic improvements and increased data availability. However, state-of-the-art models are often trained on as few as hundreds or thousands of unique object instances, and as a result generalization can be a challenge.…

Cited by 194SourceScholar
2018

Overcoming Exploration in Reinforcement Learning with Demonstrations

ICRA 2018poster

Exploration in environments with sparse rewards has been a persistent problem in reinforcement learning (RL). Many tasks are natural to specify with a sparse reward, and manually shaping a reward function can result in suboptimal performance. However, finding a non-zero reward is exponentially more…

Cited by 1038SourceScholar
2017

Hindsight Experience Replay

NeurIPS 2017poster

Dealing with sparse rewards is one of the biggest challenges in Reinforcement Learning (RL). We present a novel technique called Hindsight Experience Replay which allows sample-efficient learning from rewards which are sparse and binary and therefore avoid the need for complicated reward engineering…

Cited by 3290SourcePDFScholar