← Search

Peter Welinder

4 accepted papers

2022

Training language models to follow instructions with human feedback

NeurIPS 2022accept

Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we sho…

2018

Asymmetric Actor Critic for Image-Based Robot Learning

RSS 2018poster

Deep reinforcement learning (RL) has proven a powerful technique in many sequential decision making domains. However, robotics poses many challenges for RL, most notably training on a physical system can be expensive and dangerous, which has sparked significant interest in learning control policies…

Cited by 452SourcePDFScholar
2018

Domain Randomization and Generative Models for Robotic Grasping

IROS 2018poster

Deep learning-based robotic grasping has made significant progress thanks to algorithmic improvements and increased data availability. However, state-of-the-art models are often trained on as few as hundreds or thousands of unique object instances, and as a result generalization can be a challenge.…

Cited by 194SourceScholar
2017

Hindsight Experience Replay

NeurIPS 2017poster

Dealing with sparse rewards is one of the biggest challenges in Reinforcement Learning (RL). We present a novel technique called Hindsight Experience Replay which allows sample-efficient learning from rewards which are sparse and binary and therefore avoid the need for complicated reward engineering…

Cited by 3290SourcePDFScholar