← Search

Jan Ole von Hartz

5 accepted papers

2026

MSG: Multi-Stream Generative Policies for Sample-Efficient Robotic Manipulation

RA-L 2026

Generative robot policies such as Flow Matching offer flexible, multi-modal policy learning but are sample-inefficient. Although object-centric policies improve sample efficiency, it does not resolve this limitation. In this work, we propose Multi-Stream Generative Policy (MSG), an inference-time co

Cited by 0SourceScholar
2025

Robotic Task Ambiguity Resolution via Natural Language Interaction

IROS 2025

Language-Conditioned robotic policies allow users to specify tasks using natural language. While much research has focused on improving the action prediction of language-conditioned policies, reasoning about task descriptions has been largely overlooked. Ambiguous task descriptions often lead to dow

Cited by 5SourceScholar
2025

Whole-Body Teleoperation for Mobile Manipulation at Zero Added Cost

RA-L 2025

Demonstration data plays a key role in learning complex behaviors and training robotic foundation models. While effective control interfaces exist for static manipulators, data collection remains cumbersome and time intensive for mobile manipulators due to their large number of degrees of freedom. W

Cited by 12SourceScholar
2024

The Art of Imitation: Learning Long-Horizon Manipulation Tasks From Few Demonstrations

RA-L 2024

Task Parametrized Gaussian Mixture Models (TP-GMM) are a sample-efficient method for learning object-centric robot manipulation tasks. However, there are several open challenges to applying TP-GMMs in the wild. In this work, we tackle three crucial challenges synergistically. First, end-effector vel

Cited by 15SourcecodeScholar
2023

The Treachery of Images: Bayesian Scene Keypoints for Deep Policy Learning in Robotic Manipulation

RA-L 2023

In policy learning for robotic manipulation, sample efficiency is of paramount importance. Thus, learning and extracting more compact representations from camera observations is a promising avenue. However, current methods often assume full observability of the scene and struggle with scale invarian

Cited by 15SourcecodeScholar