← Search

Rudolf Lioutikov

32 accepted papers

2026

MoRe-ERL: Learning Motion Residuals Using Episodic Reinforcement Learning

ICRA 2026poster

We propose MoRe-ERL, a framework that combines Episodic Reinforcement Learning (ERL) and residual learning, which refines preplanned reference trajectories into safe, feasible, and efficient task-specific trajectories. This framework is general enough to incorporate into arbitrary ERL methods and mo…

2026

NaviTrace: Evaluating Embodied Navigation of Vision-Language Models

ICRA 2026poster

Vision–language models demonstrate unprecedented performance and generalization across a wide range of tasks and scenarios. Integrating these foundation models into robotic navigation systems opens pathways toward building general-purpose robots. Yet, evaluating these models’ navigation capabilities…

2026

SIR: Structured Image Representations for Explainable Robot Learning

CVPR 2026

Existing robot policies based on learned visual embeddings lack explicit structure and are sensitive to visual distractions.Thus, the representations that drive their behaviour are often opaque, making their decision-making process difficult to interpret.To address this, we introduce Structured Imag

Cited by 0SourceScholar
2025

BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning

NeurIPS 2025poster

We present the B-spline Encoded Action Sequence Tokenizer (BEAST), a novel action tokenizer that encodes action sequences into compact discrete or continuous tokens using B-splines. In contrast to existing action tokenizers based on vector quantization or byte pair encoding, BEAST requires no separ…

Cited by 0SourceScholar
2025

Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning

ICLR 2025poster

Diffusion Policies have become widely used in Imitation Learning, offering several appealing properties, such as generating multimodal and discontinuous behavior. As models are becoming larger to capture more complex capabilities, their computational demands increase, as shown by recent scaling laws…

2025

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Flow Models

CoRL 2025poster

Developing efficient Vision-Language-Action (VLA) policies is crucial for practical robotics deployment, yet current approaches face prohibitive computational costs and resource requirements. Existing diffusion-based VLA policies require multi-billion-parameter models and massive datasets to achieve…

Cited by 0SourceScholar
2025

IRIS: An Immersive Robot Interaction System

CoRL 2025poster

This paper introduces IRIS, an Immersive Robot Interaction System leveraging Extended Reality (XR). Existing XR-based systems enable efficient data collection but are often challenging to reproduce and reuse due to their specificity to particular robots, objects, simulators, and environments. IRIS a…

Cited by 0SourceScholar
2025

PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning

NeurIPS 2025poster

Robotic manipulation systems benefit from complementary sensing modalities, where each provides unique environmental information. Point clouds capture detailed geometric structure, while RGB images provide rich semantic context. Current point cloud methods struggle to capture fine-grained detail, es…

Cited by 0SourcecodeScholar
2025

TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement Learning

ICLR 2025spotlight

This work introduces Transformer-based Off-Policy Episodic Reinforcement Learning (TOP-ERL), a novel algorithm that enables off-policy updates in the ERL framework. In ERL, policies predict entire action trajectories over multiple time steps instead of single actions at every time step. These trajec…

2024

A Retrospective on the Robot Air Hockey Challenge: Benchmarking Robust, Reliable, and Safe Learning Techniques for Real-world Robotics

NeurIPS 2024poster

Machine learning methods have a groundbreaking impact in many application domains, but their application on real robotic platforms is still limited. Despite the many challenges associated with combining machine learning technology with robotics, robot learning remains one of the most promising direc…

Cited by 0SourcePDFScholar
2024

MaIL: Improving Imitation Learning with Selective State Space Models

CoRL 2024poster

This work introduces Mamba Imitation Learning (MaIL), a novel imitation learning (IL) architecture that offers a computationally efficient alternative to state-of-the-art (SoTA) Transformer policies. Transformer-based policies have achieved remarkable results due to their ability in handling human-r…

Cited by 7SourceScholar
2024

Movement Primitive Diffusion: Learning Gentle Robotic Manipulation of Deformable Objects

RA-L 2024

Policy learning in robot-assisted surgery (RAS) lacks data efficient and versatile methods that exhibit the desired motion quality for delicate surgical interventions. To this end, we introduce Movement Primitive Diffusion (MPD), a novel method for imitation learning (IL) in RAS that focuses on gent

Cited by 76SourcecodeScholar
2024

Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

RSS 2024poster

This work introduces the Multimodal Diffusion Transformer (MDT), a novel diffusion policy framework, that excels at learning versatile behavior from multimodal goal specifications with few language annotations. MDT leverages a diffusion based multimodal transformer backbone and two self-supervised a…

2024

Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning

ICLR 2024poster

Current advancements in reinforcement learning (RL) have predominantly focused on learning step-based policies that generate actions for each perceived state. While these methods efficiently leverage step information from environmental interaction, they often ignore the temporal correlation between…

2024

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

CoRL 2024poster

A central challenge towards developing robots that can relate human language to their perception and actions is the scarcity of natural language annotations in diverse robot datasets. Moreover, robot policies that follow natural language instructions are typically trained on either templated languag…

Cited by 7SourceScholar
2024

Towards Diverse Behaviors: A Benchmark for Imitation Learning with Human Demonstrations

ICLR 2024poster

Imitation learning with human data has demonstrated remarkable success in teaching robots in a wide range of skills. However, the inherent diversity in human behavior leads to the emergence of multi-modal data distributions, thereby presenting a formidable challenge for existing imitation learning a…

Cited by 24SourcePDFScholar
2024

Variational Distillation of Diffusion Policies into Mixture of Experts

NeurIPS 2024poster

This work introduces Variational Diffusion Distillation (VDD), a novel method that distills denoising diffusion policies into Mixtures of Experts (MoE) through variational inference. Diffusion Models are the current state-of-the-art in generative modeling due to their exceptional ability to accurate…

2023

Curriculum-Based Imitation of Versatile Skills

ICRA 2023poster

Learning skills by imitation is a promising concept for the intuitive teaching of robots. A common way to learn such skills is to learn a parametric model by maximizing the likelihood given the demonstrations. Yet, human demonstrations are often multi-modal, i.e., the same task is solved in multiple…

Cited by 4SourcecodeScholar
2023

Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

RSS 2023poster

We propose a new policy representation based on score-based diffusion models (SDMs). We apply our new policy representation in the domain of Goal-Conditioned Imitation Learning (GCIL) to learn general-purpose goal-specified policies from large uncurated datasets without rewards. Our new goal-conditi…

2023

Information Maximizing Curriculum: A Curriculum-Based Approach for Learning Versatile Skills

NeurIPS 2023poster

Imitation learning uses data for training policies to solve complex tasks. However, when the training data is collected from human demonstrators, it often leads to multimodal distributions because of the variability in human actions. Most imitation learning methods rely on a maximum likelihood (ML)…

Cited by 15SourcePDFScholar
2023

ProDMP: A Unified Perspective on Dynamic and Probabilistic Movement Primitives

RA-L 2023

Movement Primitives (MPs) are a well-known concept to represent and generate modular trajectories. MPs can be broadly categorized into two types: (a) dynamics-based approaches that generate smooth trajectories from any initial state, e. g., Dynamic Movement Primitives (DMPs), and (b) probabilistic a

Cited by 57SourceScholar
2022

Understanding Acoustic Patterns of Human Teachers Demonstrating Manipulation Tasks to Robots

IROS 2022poster

Humans use audio signals in the form of spoken language or verbal reactions effectively when teaching new skills or tasks to other humans. While demonstrations allow humans to teach robots in a natural way, learning from trajectories alone does not leverage other available modalities including audio…

Cited by 3SourceScholar
2021

Distributional Depth-Based Estimation of Object Articulation Models

CoRL 2021poster

We propose a method that efficiently learns distributions over articulation models directly from depth images without the need to know articulation model categories a priori. By contrast, existing methods that learn articulation models from raw observations require objects to be textured, and most o…

Cited by 26SourcecodeScholar
2021

ScrewNet: Category-Independent Articulation Model Estimation From Depth Images Using Screw Theory

ICRA 2021poster

Robots in human environments will need to interact with a wide variety of articulated objects such as cabinets, drawers, and dishwashers while assisting humans in performing day-to-day tasks. Existing methods either require objects to be textured or need to know the articulation model category a pri…

Cited by 96SourcecodeScholar
2021

Self-Supervised Online Reward Shaping in Sparse-Reward Environments

IROS 2021poster

We introduce Self-supervised Online Reward Shaping (SORS), which aims to improve the sample efficiency of any RL algorithm in sparse-reward environments by automatically densifying rewards. The proposed framework alternates between classification-based reward inference and policy update steps—the or…

Cited by 66SourcecodeScholar
2018

Inducing Probabilistic Context-Free Grammars for the Sequencing of Movement Primitives

ICRA 2018poster

Movement Primitives are a well studied and widely applied concept in modern robotics. Composing primitives out of an existing library, however, has shown to be a challenging problem. We propose the use of probabilistic context-free grammars to sequence a series of primitives to generate complex robo…

Cited by 11SourceScholar
2017

Guiding Trajectory Optimization by Demonstrated Distributions

RA-L 2017

Trajectory optimization is an essential tool for motion planning under multiple constraints of robotic manipulators. Optimization-based methods can explicitly optimize a trajectory by leveraging prior knowledge of the system and have been used in various applications such as collision avoidance. How

Cited by 60SourceScholar
2015

Learning multiple collaborative tasks with a mixture of Interaction Primitives

ICRA 2015poster

Robots that interact with humans must learn to not only adapt to different human partners but also to new interactions. Such a form of learning can be achieved by demonstrations and imitation. A recently introduced method to learn interactions from demonstrations is the framework of Interaction Prim…

Cited by 145SourceScholar
2015

Model-Based Relative Entropy Stochastic Search

NeurIPS 2015poster

Stochastic search algorithms are general black-box optimizers. Due to their ease of use and their generality, they have recently also gained a lot of attention in operations research, machine learning and policy search. Yet, these algorithms require a lot of evaluations of the objective, scale poorl…

Cited by 106SourcePDFScholar