← Search

Adrien Ecoffet

5 accepted papers

2024

Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

ICML 2024oral

Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior---for example, to evaluate whether a model faithfully followed instructions or generated safe outputs. However, future superhuman models will behave…

Cited by 260SourcePDFScholar
2022

Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos

NeurIPS 2022accept

Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision domains such as robotics, video games, and computer use, publicly available data do…

2020

Estimating Q(s,s’) with Deep Deterministic Dynamics Gradients

ICML 2020poster

In this paper, we introduce a novel form of value function, $Q(s, s’)$, that expresses the utility of transitioning from a state $s$ to a neighboring state $s’$ and then acting optimally thereafter. In order to derive an optimal policy, we develop a forward dynamics model that learns to make next-st…

Cited by 26SourcePDFScholar
2020

Exploration Based Language Learning for Text-Based Games

IJCAI 2020poster

This work presents an exploration and imitation-learning-based agent capable of state-of-the-art performance in playing text-based computer games. These games are of interest as they can be seen as a testbed for language understanding, problem-solving, and language generation by artificial agents.…

Cited by 0SourcePDFScholar