← Search

Carlos Florensa

8 accepted papers

2021

Which Mutual-Information Representation Learning Objectives are Sufficient for Control?

NeurIPS 2021poster

Mutual information (MI) maximization provides an appealing formalism for learning representations of data. In the context of reinforcement learning (RL), such representations can accelerate learning by discarding irrelevant and redundant information, while retaining the information necessary for con…

Cited by 41SourcePDFScholar
2020

Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning

ICRA 2020poster

Traditional robotic approaches rely on an accurate model of the environment, a detailed description of how to perform the task, and a robust perception system to keep track of the current state. On the other hand, reinforcement learning approaches can operate directly from raw sensory inputs with on…

Cited by 75SourceScholar
2020

Sub-policy Adaptation for Hierarchical Reinforcement Learning

ICLR 2020poster

Hierarchical reinforcement learning is a promising approach to tackle long-horizon decision-making problems with sparse rewards. Unfortunately, most methods still decouple the lower-level skill acquisition process and the training of a higher level that controls the skills in a new task. Leaving the…

Cited by 106SourceScholar
2018

Automatic Goal Generation for Reinforcement Learning Agents

ICML 2018oral

Reinforcement learning (RL) is a powerful technique to train an agent to perform a task; however, an agent that is trained using RL is only capable of achieving the single task that is specified via its reward function. Such an approach does not scale well to settings in which an agent needs to perf…

Cited by 530SourcePDFScholar
2017

Reverse Curriculum Generation for Reinforcement Learning

CoRL 2017

Many relevant tasks require an agent to reach a certain state, or to manipulate objects into a desired configuration. For example, we might want a robot to align and assemble a gear onto an axle or insert and turn a key in a lock. These goal-oriented tasks present a considerable challenge for reinfo

Cited by 0SourcePDFScholar