← Search

Moritz Reuss

11 accepted papers

2026

NaviTrace: Evaluating Embodied Navigation of Vision-Language Models

ICRA 2026poster

Vision–language models demonstrate unprecedented performance and generalization across a wide range of tasks and scenarios. Integrating these foundation models into robotic navigation systems opens pathways toward building general-purpose robots. Yet, evaluating these models’ navigation capabilities…

2025

BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning

NeurIPS 2025poster

We present the B-spline Encoded Action Sequence Tokenizer (BEAST), a novel action tokenizer that encodes action sequences into compact discrete or continuous tokens using B-splines. In contrast to existing action tokenizers based on vector quantization or byte pair encoding, BEAST requires no separ…

Cited by 0SourceScholar
2025

Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning

ICLR 2025poster

Diffusion Policies have become widely used in Imitation Learning, offering several appealing properties, such as generating multimodal and discontinuous behavior. As models are becoming larger to capture more complex capabilities, their computational demands increase, as shown by recent scaling laws…

2025

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Flow Models

CoRL 2025poster

Developing efficient Vision-Language-Action (VLA) policies is crucial for practical robotics deployment, yet current approaches face prohibitive computational costs and resource requirements. Existing diffusion-based VLA policies require multi-billion-parameter models and massive datasets to achieve…

Cited by 0SourceScholar
2025

PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning

NeurIPS 2025poster

Robotic manipulation systems benefit from complementary sensing modalities, where each provides unique environmental information. Point clouds capture detailed geometric structure, while RGB images provide rich semantic context. Current point cloud methods struggle to capture fine-grained detail, es…

Cited by 0SourcecodeScholar
2024

Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

RSS 2024poster

This work introduces the Multimodal Diffusion Transformer (MDT), a novel diffusion policy framework, that excels at learning versatile behavior from multimodal goal specifications with few language annotations. MDT leverages a diffusion based multimodal transformer backbone and two self-supervised a…

2024

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

CoRL 2024poster

A central challenge towards developing robots that can relate human language to their perception and actions is the scarcity of natural language annotations in diverse robot datasets. Moreover, robot policies that follow natural language instructions are typically trained on either templated languag…

Cited by 7SourceScholar
2024

Towards Diverse Behaviors: A Benchmark for Imitation Learning with Human Demonstrations

ICLR 2024poster

Imitation learning with human data has demonstrated remarkable success in teaching robots in a wide range of skills. However, the inherent diversity in human behavior leads to the emergence of multi-modal data distributions, thereby presenting a formidable challenge for existing imitation learning a…

Cited by 24SourcePDFScholar
2023

Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

RSS 2023poster

We propose a new policy representation based on score-based diffusion models (SDMs). We apply our new policy representation in the domain of Goal-Conditioned Imitation Learning (GCIL) to learn general-purpose goal-specified policies from large uncurated datasets without rewards. Our new goal-conditi…

2023

Information Maximizing Curriculum: A Curriculum-Based Approach for Learning Versatile Skills

NeurIPS 2023poster

Imitation learning uses data for training policies to solve complex tasks. However, when the training data is collected from human demonstrators, it often leads to multimodal distributions because of the variability in human actions. Most imitation learning methods rely on a maximum likelihood (ML)…

Cited by 15SourcePDFScholar
2022

End-to-End Learning of Hybrid Inverse Dynamics Models for Precise and Compliant Impedance Control

RSS 2022poster

It is well-known that inverse dynamics models can improve tracking performance in robot control. These models need to precisely capture the robot dynamics, which consist of well-understood components, e.g., rigid body dynamics, and effects that remain challenging to capture, e.g., stick-slip frictio…

Cited by 11SourcePDFScholar