← Search

Stefan Schaal

60 accepted papers

2026

The Trajectory Bundle Method: Unifying Sequential-Convex Programming and Sampling-Based Trajectory Optimization

ICRA 2026poster

We present a unified framework for solving trajectory optimization problems in a derivative-free manner through the use of sequential convex programming. Traditionally, nonconvex optimization problems are solved by forming and solving a sequence of convex optimization problems, where the cost and co…

2025

Efficient Online Learning of Contact Force Models for Connector Insertion

ICRA 2025

Contact-rich manipulation tasks with stiff frictional elements, like connector insertion, are difficult to model with rigid-body simulators. In this work, we propose a new approach for modeling these environments by learning a quasistatic contact force model instead of a full simulator. Using a feat

Cited by 6SourcecodeScholar
2024

A Comparison of Imitation Learning Algorithms for Bimanual Manipulation

RA-L 2024

Amidst the wide popularity of imitation learning algorithms in robotics, their properties regarding hyperparameter sensitivity, ease of training, data efficiency, and performance have not been well-studied in high-precision industry-inspired environments. In this work, we demonstrate the limitations

Cited by 22SourceScholar
2024

GenCHiP: Generating Robot Policy Code for High-Precision and Contact-Rich Manipulation Tasks

IROS 2024poster

Large Language Models (LLMs) have been successful at generating robot policy code, but so far these results have been limited to high-level tasks that do not require precise movement. It is an open question how well such approaches work for tasks that require reasoning over contact forces and workin…

Cited by 5SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

RT-Sketch: Goal-Conditioned Imitation Learning from Hand-Drawn Sketches

CoRL 2024poster

Natural language and images are commonly used as goal representations in goal-conditioned imitation learning. However, language can be ambiguous and images can be over-specified. In this work, we study hand-drawn sketches as a modality for goal specification. Sketches can be easy to provide on the f…

Cited by 11SourcecodeScholar
2024

SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning

ICRA 2024poster

In recent years, significant progress has been made in the field of robotic reinforcement learning (RL), enabling methods that handle complex image observations, train in the real world, and incorporate auxiliary data, such as demonstrations and prior experience. However, despite these advances, rob…

Cited by 48SourcecodeScholar
2023

Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning

ICRA 2023poster

Humans are capable of abstracting various tasks as different combinations of multiple attributes. This perspective of compositionality is vital for human rapid learning and adaption since previous experiences from related tasks can be combined to generalize across novel compositional settings. In th…

Cited by 14SourceScholar
2022

A System for Imitation Learning of Contact-Rich Bimanual Manipulation Policies

IROS 2022poster

In this paper, we discuss a framework for teaching bimanual manipulation tasks by imitation. To this end, we present a system and algorithms for learning compliant and contact-rich robot behavior from human demonstrations. The presented system combines insights from admittance control and machine le…

Cited by 34SourceScholar
2022

CaTGrasp: Learning Category-Level Task-Relevant Grasping in Clutter from Simulation

ICRA 2022poster

Task-relevant grasping is critical for industrial assembly, where downstream manipulation tasks constrain the set of valid grasps. Learning how to perform this task, however, is challenging, since task-relevant grasp labels are hard to define and annotate. There is also yet no consensus on proper re…

Cited by 99SourcecodeScholar
2022

Efficient Spatial Representation and Routing of Deformable One-Dimensional Objects for Manipulation

IROS 2022poster

With the field of rigid-body robotics having matured in the last fifty years, routing, planning, and manipulation of deformable objects have recently emerged as a more untouched research area in many fields ranging from surgical robotics to industrial assembly and construction. Routing approaches fo…

Cited by 11SourceScholar
2022

Multi-Task Learning with Sequence-Conditioned Transporter Networks

ICRA 2022poster

Enabling robots to solve multiple manipulation tasks has a wide range of industrial applications. While learning-based approaches enjoy flexibility and generalizability, scaling these approaches to solve such compositional tasks remains a challenge. In this work, we aim to solve multi-task learning…

Cited by 15SourceScholar
2022

Offline Meta-Reinforcement Learning for Industrial Insertion

ICRA 2022poster

Reinforcement learning (RL) can in principle let robots automatically adapt to new tasks, but current RL methods require a large number of trials to accomplish this. In this paper, we tackle rapid adaptation to new tasks through the framework of meta-learning, which utilizes past tasks to learn to a…

Cited by 104SourceScholar
2022

Residual Learning From Demonstration: Adapting DMPs for Contact-Rich Manipulation

RA-L 2022

Manipulation skills involving contact and friction are inherent to many robotics tasks. Using the class of motor primitives for peg-in-hole like insertions, we study how robots can learn such skills. Dynamic Movement Primitives (DMP) are a popular way of extracting such policies through behaviour cl

Cited by 67SourceScholar
2022

Symbolic State Estimation with Predicates for Contact-Rich Manipulation Tasks

ICRA 2022poster

Manipulation tasks often require a robot to adjust its sensorimotor skills based on the state it finds itself in. Taking peg-in-hole as an example: once the peg is aligned with the hole, the robot should push the peg downwards. While high level execution frameworks such as state machines and behavio…

Cited by 13SourceScholar
2022

Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation

ICLR 2022poster

Complex sequential tasks in continuous-control settings often require agents to successfully traverse a set of ``narrow passages'' in their state space. Solving such tasks with a sparse reward in a sample-efficient manner poses a challenge to modern reinforcement learning (RL) due to the associated…

Cited by 19SourcePDFScholar
2022

You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

RSS 2022poster

Promising results have been achieved recently in category-level manipulation that generalizes across object instances. Nevertheless, it often requires expensive real-world data collection and manual specification of semantic keypoints for each object category and task. Additionally, coarse keypoint…

2021

Benchmarking Off-The-Shelf Solutions to Robotic Assembly Tasks

IROS 2021poster

In recent years, many learning based approaches have been studied to realize robotic manipulation and assembly tasks, often including vision and force/tactile feedback. How-ever, it is unclear what the baseline state-of-the-art performance is and what the bottleneck problems are. In this work, we ev…

Cited by 28SourceScholar
2021

Learning Dense Rewards for Contact-Rich Manipulation Tasks

ICRA 2021poster

Rewards play a crucial role in reinforcement learning. To arrive at the desired policy, the design of a suitable reward function often requires significant domain expertise as well as trial-and-error. Here, we aim to minimize the effort involved in designing reward functions for contact-rich manipul…

Cited by 50SourceScholar
2019

A Robustness Analysis of Inverse Optimal Control of Bipedal Walking

RA-L 2019

Cost functions have the potential to provide compact and understandable generalizations of motion. The goal of inverse optimal control (IOC) is to analyze an observed behavior which is assumed to be optimal with respect to an unknown cost function, and infer this cost function. Here we develop a met

Cited by 11SourceScholar
2018

Learning Manipulation Graphs from Demonstrations Using Multimodal Sensory Signals

ICRA 2018poster

Complex contact manipulation tasks can be decomposed into sequences of motor primitives. Individual primitives often end with a distinct contact state, such as inserting a screwdriver tip into a screw head or loosening it through twisting. To achieve robust execution, the robot should be able to ver…

Cited by 35SourceScholar
2018

Learning Sensor Feedback Models from Demonstrations via Phase-Modulated Neural Networks

ICRA 2018poster

In order to robustly execute a task under environmental uncertainty, a robot needs to be able to reactively adapt to changes arising in its environment. The environment changes are usually reflected in deviation from expected sensory traces. These deviations in sensory traces can be used to drive th…

Cited by 24SourceScholar
2018

On Time Optimization of Centroidal Momentum Dynamics

ICRA 2018poster

Recently, the centroidal momentum dynamics has received substantial attention to plan dynamically consistent motions for robots with arms and legs in multi-contact scenarios. However, it is also non convex which renders any optimization approach difficult and timing is usually kept fixed in most tra…

Cited by 77SourceScholar
2018

Probabilistic Recurrent State-Space Models

ICML 2018oral

State-space models (SSMs) are a highly expressive model class for learning patterns in time series data and for system identification. Deterministic versions of SSMs (e.g., LSTMs) proved extremely successful in modeling complex time series data. Fully probabilistic SSMs, however, are often found har…

2018

Real-Time Perception Meets Reactive Motion Generation

RA-L 2018

We address the challenging problem of robotic grasping and manipulation in the presence of uncertainty. This uncertainty is due to noisy sensing, inaccurate models, and hard-to-predict environment dynamics. We quantify the importance of continuous, real-time perception and its tight integration with

Cited by 120SourceScholar
2018

Time-Contrastive Networks: Self-Supervised Learning from Video

ICRA 2018poster

We propose a self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this representation can be used in two robotic imitation settings: imitating object interactions from videos of humans, and imitati…

Cited by 1009SourcecodeScholar
2017

Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning

ICML 2017poster

Reinforcement learning algorithms for real-world robotic applications must be able to handle complex, unknown dynamical systems while maintaining data-efficient learning. These requirements are handled well by model-free and model-based RL approaches, respectively. In this work, we aim to combine th…

Cited by 227SourcePDFScholar
2017

Learning feedback terms for reactive planning and control

ICRA 2017poster

With the advancement of robotics, machine learning, and machine perception, increasingly more robots will enter human environments to assist with daily tasks. However, dynamically-changing human environments requires reactive motion plans. Reactivity can be accomplished through re-planning, e.g. mod…

Cited by 56SourceScholar
2017

Model-based policy search for automatic tuning of multivariate PID controllers

ICRA 2017poster

PID control architectures are widely used in industrial applications. Despite their low number of open parameters, tuning multiple, coupled PID controllers can become tedious in practice. In this paper, we extend PILCO, a model-based policy search framework, to automatically tune multivariate PID co…

Cited by 49SourceScholar
2017

Multi-Modal Imitation Learning from Unstructured Demonstrations using Generative Adversarial Nets

NeurIPS 2017poster

Imitation learning has traditionally been applied to learn a single task from demonstrations thereof. The requirement of structured and isolated demonstrations limits the scalability of imitation learning approaches as they are difficult to apply to real-world scenarios, where robots have to be able…

Cited by 202SourcePDFScholar
2017

On the relevance of grasp metrics for predicting grasp success

IROS 2017poster

We aim to reliably predict whether a grasp on a known object is successful before it is executed in the real world. There is an entire suite of grasp metrics that has already been developed which rely on precisely known contact points between object and hand. However, it remains unclear whether and…

Cited by 38SourceScholar
2017

Optimizing Long-term Predictions for Model-based Policy Search

CoRL 2017

We propose a novel long-term optimization criterion to improve the robustness of model-based reinforcement learning in real-world scenarios. Learning a dynamics model to derive a solution promises much greater data-efficiency and reusability compared to model-free alternatives. In practice, however,

2017

Path integral guided policy search

ICRA 2017poster

We present a policy search method for learning complex feedback control policies that map from high-dimensional sensory inputs to motor torques, for manipulation tasks with discontinuous contact dynamics. We build on a prior technique called guided policy search (GPS), which iteratively optimizes a…

Cited by 208SourceScholar
2017

Probabilistic Articulated Real-Time Tracking for Robot Manipulation

RA-L 2017

We propose a probabilistic filtering method which fuses joint measurements with depth images to yield a precise, real-time estimate of the end-effector pose in the camera frame. This avoids the need for frame transformations when using it in combination with visual object tracking methods. Precision

Cited by 75SourceScholar
2017

Virtual vs. real: Trading off simulations and physical experiments in reinforcement learning with Bayesian optimization

ICRA 2017poster

In practice, the parameters of control policies are often tuned manually. This is time-consuming and frustrating. Reinforcement learning is a promising alternative that aims to automate this process, yet often requires too many experiments to be practical. In this paper, we propose a solution to thi…

Cited by 176SourceScholar
2016

Automatic LQR tuning based on Gaussian process global optimization

ICRA 2016poster

This paper proposes an automatic controller tuning framework based on linear optimal control combined with Bayesian optimization. With this framework, an initial set of controller gains is automatically improved according to a pre-defined performance objective evaluated from experimental data. The u…

Cited by 219SourceScholar
2016

Depth-based object tracking using a Robust Gaussian Filter

ICRA 2016

We consider the problem of model-based 3D-tracking of objects given dense depth images as input. Two difficulties preclude the application of a standard Gaussian filter to this problem. First of all, depth sensors are characterized by fat-tailed measurement noise. To address this issue, we show how

Cited by 86SourceScholar
2016

Learning where to search using visual attention

IROS 2016poster

One of the central tasks for a household robot is searching for specific objects. It does not only require localizing the target object but also identifying promising search locations in the scene if the target is not immediately visible. As computation time and hardware resources are usually limite…

Cited by 3SourceScholar
2016

Robot arm pose estimation by pixel-wise regression of joint angles

ICRA 2016

To achieve accurate vision-based control with a robotic arm, a good hand-eye coordination is required. However, knowing the current configuration of the arm can be very difficult due to noisy readings from joint encoders or an inaccurate hand-eye calibration. We propose an approach for robot arm pos

Cited by 42SourceScholar
2016

Self-supervised regrasping using spatio-temporal tactile features and reinforcement learning

IROS 2016poster

We introduce a framework for learning regrasping behaviors based on tactile data. First, we present a grasp stability predictor that uses spatio-temporal tactile features collected from the early-object-lifting phase to predict the grasp outcome with a high accuracy. Next, the trained predictor is u…

Cited by 105SourceScholar
2016

Structured contact force optimization for kino-dynamic motion generation

IROS 2016poster

Optimal control approaches in combination with trajectory optimization have recently proven to be a promising control strategy for legged robots. Computationally efficient and robust algorithms were derived using simplified models of the contact interaction between robot and environment such as the…

Cited by 99SourceScholar
2016

Warping the workspace geometry with electric potentials for motion optimization of manipulation tasks

IROS 2016poster

In this paper we present motion optimization algorithms for computing manipulation motions in presence of obstacles. Our approach builds a geometric representation of the workspace by constructing Riemannian metrics using electric potentials emanating from the workspace obstacles. Velocity of the ro…

Cited by 20SourceScholar
2015

A New Perspective and Extension of the Gaussian Filter

RSS 2015poster

The Gaussian Filter (GF) is one of the most widely used filtering algorithms; instances are the Extended Kalman Filter, the Unscented Kalman Filter and the Divided Difference Filter. GFs represent the belief of the current state by a Gaussian with the mean being an affine function of the measurement…

Cited by 32SourcePDFScholar
2015

Data-Driven Online Decision Making for Autonomous Manipulation

RSS 2015poster

One of the main challenges in autonomous manipulation is to generate appropriate multi-modal reference trajectories that enable feedback controllers to compute control commands that compensate for unmodeled perturbations and therefore to achieve the task at hand. We propose a data-driven approach to…

Cited by 79SourcePDFScholar
2015

Direct Loss Minimization Inverse Optimal Control

RSS 2015poster

Inverse Optimal Control (IOC) has strongly impacted the systems engineering process, enabling automated planner tuning through straightforward and intuitive demonstration. The most successful and established applications, though, have been in lower dimensional problems such as navigation planning wh…

Cited by 57SourcePDFScholar
2015

The Coordinate Particle Filter - a novel Particle Filter for high dimensional systems

ICRA 2015poster

Parametric filters, such as the Extended Kalman Filter and the Unscented Kalman Filter, typically scale well with the dimensionality of the problem, but they are known to fail if the posterior state distribution cannot be closely approximated by a density of the assumed parametric form.

Cited by 12SourceScholar