← Search

Sebastian Blaes

9 accepted papers

2023

Benchmarking Offline Reinforcement Learning on Real-Robot Hardware

ICLR 2023top-25%

Learning policies from previously recorded data is a promising direction for real-world robotics tasks, as online learning is often infeasible. Dexterous manipulation in particular remains an open problem in its general form. The combination of offline reinforcement learning with large diverse datas…

Cited by 38SourcePDFScholar
2023

Optimistic Active Exploration of Dynamical Systems

NeurIPS 2023poster

Reinforcement learning algorithms commonly seek to optimize policies for solving one particular task. How should we explore an unknown dynamical system such that the estimated model allows us to solve multiple downstream tasks in a zero-shot manner? In this paper, we address this challenge, by deve…

Cited by 12SourcePDFScholar
2023

Versatile Skill Control via Self-supervised Adversarial Imitation of Unlabeled Mixed Motions

ICRA 2023poster

Learning diverse skills is one of the main challenges in robotics. To this end, imitation learning approaches have achieved impressive results. These methods require explicitly labeled datasets or assume consistent skill execution to enable learning and active control of individual behaviors, which…

Cited by 34SourceScholar
2022

Curious Exploration via Structured World Models Yields Zero-Shot Object Manipulation

NeurIPS 2022accept

It has been a long-standing dream to design artificial agents that explore their environment efficiently via intrinsic motivation, similar to how children perform curious free play. Despite recent advances in intrinsically motivated reinforcement learning (RL), sample-efficient exploration in object…

Cited by 30SourcePDFScholar
2022

Learning Agile Skills via Adversarial Imitation of Rough Partial Demonstrations

CoRL 2022oral

Learning agile skills is one of the main challenges in robotics. To this end, reinforcement learning approaches have achieved impressive results. These methods require explicit task information in terms of a reward function or an expert that can be queried in simulation to provide a target control o…

Cited by 73SourceScholar
2021

Extracting Strong Policies for Robotics Tasks from Zero-Order Trajectory Optimizers

ICLR 2021poster

Solving high-dimensional, continuous robotic tasks is a challenging optimization problem. Model-based methods that rely on zero-order optimizers like the cross-entropy method (CEM) have so far shown strong performance and are considered state-of-the-art in the model-based reinforcement learning comm…

Cited by 13SourcePDFScholar
2020

Sample-efficient Cross-Entropy Method for Real-time Planning

CoRL 2020

Trajectory optimizers for model-based reinforcement learning, such as the Cross-Entropy Method (CEM), can yield compelling results even in high-dimensional control tasks and sparse-reward environments. However, their sampling inefficiency prevents them from being used for real-time planning and cont

2019

Control What You Can: Intrinsically Motivated Task-Planning Agent

NeurIPS 2019poster

We present a novel intrinsically motivated agent that learns how to control the environment in a sample efficient manner, that is with as few environment interactions as possible, by optimizing learning progress. It learns what can be controlled, how to allocate time and attention as well as the rel…