← Search

Sören Schwertfeger

17 accepted papers

2026

From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings

CVPR 2026

We present a novel unsupervised framework to unlock vast unlabeled human demonstration data from continuous industrial video streams for Vision-Language-Action (VLA) model pre-training. Our method first trains a lightweight motion tokenizer to encode motion dynamics, then employs an unsupervised act

Cited by 0SourcecodeScholar
2026

OsmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation Via Semantic Maps and Large Language Models Reasoning

ICRA 2026poster

Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features, achieving a high level of detail and guiding robots to find objects specified by open-vocabulary language queries. While the issue of scalability for such approaches has received some …

2026

osmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation via Semantic Maps and Large Language Models Reasoning

RA-L 2026

Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features, achieving a high level of detail and guiding robots to find objects specified by open-vocabulary language queries. While the issue of scalability for such approaches has received some

Cited by 5SourcecodeScholar
2025

Adversarial Locomotion and Motion Imitation for Humanoid Policy Learning

NeurIPS 2025poster

Humans exhibit diverse and expressive whole-body movements. However, attaining human-like whole-body coordination in humanoid robots remains challenging, as conventional approaches that mimic whole-body motions often neglect the distinct roles of upper and lower body. This oversight leads to computa…

Cited by 0SourcecodeScholar
2025

Intelligent LiDAR Navigation: Leveraging External Information and Semantic Maps with LLM as Copilot

IROS 2025

Traditional robot navigation systems primarily utilize occupancy grid maps and laser-based sensing technologies, as demonstrated by the popular move_base package in ROS. Unlike robots, humans navigate not only through spatial awareness and physical distances but also by integrating external informat

Cited by 4SourcecodeScholar
2024

RealDex: Towards Human-like Grasping for Robotic Dexterous Hand

IJCAI 2024poster

In this paper, we introduce RealDex, a pioneering dataset capturing authentic dexterous hand grasping motions infused with human behavioral patterns, enriched by multi-view and multimodal visual data. Utilizing a teleoperation system, we seamlessly synchronize human-robot hand poses in real time. Th…

2023

FloorplanNet: Learning Topometric Floorplan Matching for Robot Localization

ICRA 2023poster

Given a building floorplan, humans can localize themselves by matching the observation of the environment with the floorplan using geometric, semantic, and topological clues. Inspired by this insight, this paper proposes a learning- based topometric robot localization method FloorplanNet, which impl…

Cited by 9SourcecodeScholar
2023

Optimizing the Extended Fourier Mellin Transformation Algorithm

IROS 2023poster

With the increasing application of robots, stable and efficient Visual Odometry (VO) algorithms are becoming more and more important. Based on the Fourier Mellin Transformation (FMT) algorithm, the extended Fourier Mellin Transformation (eFMT) is an image registration approach that can be applied to…

Cited by 0SourcecodeScholar
2023

Robot Parkour Learning

CoRL 2023oral

Parkour is a grand challenge for legged locomotion that requires robots to overcome various obstacles rapidly in complex environments. Existing methods can generate either diverse but blind locomotion skills or vision-based but specialized skills by using reference animal data or complex rewards. Ho…

Cited by 195SourcecodeScholar
2022

Accurate Calibration of Multi-Perspective Cameras from a Generalization of the Hand-Eye Constraint

ICRA 2022poster

Multi-perspective cameras are quickly gaining importance in many applications such as smart vehicles and virtual or augmented reality. However, a large system size or absence of overlap in neighbouring fields-of-view often complicate their calibration. We present a novel solution which relies on the…

Cited by 10SourcecodeScholar
2022

Multical: Spatiotemporal Calibration for Multiple IMUs, Cameras and LiDARs

IROS 2022poster

Spatiotemporal calibration of sensors, especially of those which do not share their fields of view, is becoming increasingly important in the fields of autonomous driving and robotics. This paper presents a general sensor calibration method, named Multical, that makes use of multiple planar calibrat…

Cited by 12SourceScholar
2019

Adaptive Navigation Scheme for Optimal Deep-Sea Localization Using Multimodal Perception Cues

IROS 2019poster

Underwater robot interventions require a high level of safety and reliability. A major challenge to address is a robust and accurate acquisition of localization estimates, as it is a prerequisite to enable more complex tasks, e.g. floating manipulation and mapping. State-of-the-art navigation in com…

Cited by 17SourceScholar
2019

Pose Estimation for Omni-directional Cameras using Sinusoid Fitting

IROS 2019poster

We propose a novel pose estimation method for geometric vision of omni-directional cameras. On the basis of the regularity of the pixel movement after camera pose changes, we formulate and prove the sinusoidal relationship between pixels movement and camera motion. We use the improved Fourier-Mellin…

Cited by 7SourceScholar