← Search

Sven Behnke

72 accepted papers

2026

Iterative Motion Compensation for Canonical 3D Reconstruction From UAV Plant Images Captured in Windy Conditions

RA-L 2026

Three-dimensional (3D) phenotyping of plants plays a crucial role for understanding plant growth, yield prediction, and disease control. We present a pipeline capable of generating high-quality 3D reconstructions of individual agricultural plants. To acquire data, a small commercially available Unma

Cited by 0SourceScholar
2026

Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking

ICRA 2026poster

Vision-language models (VLMs) excel in visual understanding but often lack reliable grounding capabilities and actionable inference rates. Integrating them with open-vocabulary object detection (OVD), instance segmentation, and tracking leverages their strengths while mitigating these drawbacks. We …

2026

Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation

AAAI 2026technical

Cross-lingual, cross-task transfer is challenged by task-specific data scarcity, which becomes more severe as language support grows and is further amplified in vision-language models (VLMs). We investigate multilingual generalization in encoder-decoder transformer VLMs to enable zero-shot image cap

Cited by 0SourcePDFScholar
2025

Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking

RA-L 2025

Vision-language models (VLMs) excel in visual understanding but often lack reliable grounding capabilities and actionable inference rates. Integrating them with open-vocabulary object detection (OVD), instance segmentation, and tracking leverages their strengths while mitigating these drawbacks. We

Cited by 2SourceScholar
2025

PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning

ICML 2025poster

Predicting future scene representations is a crucial task for enabling robots to understand and interact with the environment. However, most existing methods rely on videos and simulations with precise action annotations, limiting their ability to leverage the large amount of avail- able unlabeled v…

Cited by 1SourcePDFScholar
2025

SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from Pixels

ICML 2025poster

Learning a latent dynamics model provides a task-agnostic representation of an agent's understanding of its environment. Leveraging this knowledge for model-based reinforcement learning (RL) holds the potential to improve sample efficiency over model-free methods by learning from imagined rollouts.…

2024

FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance Head-pose and Facial Expression Features

CVPR 2024poster

The task of face reenactment is to transfer the head motion and facial expressions from a driving video to the appearance of a source image which may be of a different person (cross-reenactment). Most existing methods are CNN-based and estimate optical flow from the source image to the current drivi…

Cited by 7SourcePDFScholar
2024

Grasp Anything: Combining Teacher-Augmented Policy Gradient Learning with Instance Segmentation to Grasp Arbitrary Objects

ICRA 2024poster

Interactive grasping from clutter, akin to human dexterity, is one of the longest-standing problems in robot learning. Challenges stem from the intricacies of visual perception, the demand for precise motor skills, and the complex interplay between the two. In this work, we present Teacher-Augmented…

Cited by 8SourcecodeScholar
2024

HortiBot: An Adaptive Multi-Arm System for Robotic Horticulture of Sweet Peppers

IROS 2024poster

Horticultural tasks such as pruning and selective harvesting are labor intensive and horticultural staff are hard to find. Automating these tasks is challenging due to the semi-structured greenhouse workspaces, changing environmental conditions such as lighting, dense plant growth with many occlusio…

Cited by 8SourceScholar
2024

MOTPose: Multi-object 6D Pose Estimation for Dynamic Video Sequences using Attention-based Temporal Fusion

ICRA 2024poster

Cluttered bin-picking environments are challenging for pose estimation models. Despite the impressive progress enabled by deep learning, single-view RGB pose estimation models perform poorly in cluttered dynamic environments. Imbuing the rich temporal information contained in the video of scenes has…

Cited by 0SourceScholar
2024

SLCF-Net: Sequential LiDAR-Camera Fusion for Semantic Scene Completion using a 3D Recurrent U-Net

ICRA 2024poster

We introduce SLCF-Net, a novel approach for the Semantic Scene Completion (SSC) task that sequentially fuses LiDAR and camera data. It jointly estimates missing geometry and semantics in a scene from sequences of RGB images and sparse LiDAR measurements. The images are semantically segmented by a pr…

Cited by 6SourcecodeScholar
2023

Attention-Based VR Facial Animation with Visual Mouth Camera Guidance for Immersive Telepresence Avatars

IROS 2023poster

Facial animation in virtual reality environments is essential for applications that necessitate clear visibility of the user's face and the ability to convey emotional signals. In our scenario, we animate the face of an operator who controls a robotic Avatar system. The use of facial animation is pa…

Cited by 4SourceScholar
2023

External Camera-Based Mobile Robot Pose Estimation for Collaborative Perception with Smart Edge Sensors

ICRA 2023poster

We present an approach for estimating a mobile robot's pose w.r.t. the allocentric coordinates of a network of static cameras using multi-view RGB images. The images are processed online, locally on smart edge sensors by deep neural networks to detect the robot and estimate 2D keypoints defined at d…

Cited by 17SourceScholar
2023

PermutoSDF: Fast Multi-View Reconstruction With Implicit Surfaces Using Permutohedral Lattices

CVPR 2023poster

Neural radiance-density field methods have become increasingly popular for the task of novel-view rendering. Their recent extension to hash-based positional encoding ensures fast training and inference with visually pleasing results. However, density-based methods struggle with recovering accurate s…

2023

Quadrupedal Footstep Planning Using Learned Motion Models of a Black-Box Controller

IROS 2023poster

Legged robots are increasingly entering new domains and applications, including search and rescue, inspection, and logistics. However, for such a systems to be valuable in real-world scenarios, they must be able to autonomously and robustly navigate irregular terrains. In many cases, robots that are…

Cited by 1SourceScholar
2022

Abstract Flow for Temporal Semantic Segmentation on the Permutohedral Lattice

ICRA 2022poster

Semantic segmentation is a core ability required by autonomous agents, as being able to distinguish which parts of the scene belong to which object class is crucial for navigation and interaction with the environment. Approaches which use only one time-step of data cannot distinguish between moving…

Cited by 18SourcecodeScholar
2022

FaDIV-Syn: Fast Depth-Independent View Synthesis using Soft Masks and Implicit Blending

RSS 2022poster

Novel view synthesis is required in many robotic applications, such as VR teleoperation and scene reconstruction. Existing methods are often too slow for these contexts, cannot handle dynamic scenes, and are limited by their explicit depth estimation stage, where incorrect depth predictions can lead…

2022

Neural Strands: Learning Hair Geometry and Appearance from Multi-View Images

ECCV 2022poster

"We present Neural Strands, a novel learning framework for modeling accurate hair geometry and appearance from multi-view image inputs. The learned hair model can be rendered in real-time from any viewpoint with high-fidelity view-dependent effects. Our model achieves intuitive shape and style contr…

Cited by 44SourcePDFScholar
2022

Real-Robot Deep Reinforcement Learning: Improving Trajectory Tracking of Flexible-Joint Manipulator with Reference Correction

ICRA 2022poster

Flexible-joint manipulators are governed by complex nonlinear dynamics, defining a challenging control problem. In this work, we propose an approach to learn an outer-loop joint trajectory tracking controller with deep reinforcement learning. The controller represented by a stochastic policy is lear…

Cited by 11SourceScholar
2021

NimbRo Avatar: Interactive Immersive Telepresence with Force-Feedback Telemanipulation

IROS 2021poster

Robotic avatars promise immersive teleoperation with human-like manipulation and communication capabilities. We present such an avatar system, based on the key components of immersive 3D visualization and transparent force-feedback telemanipulation. Our avatar robot features an anthropomorphic biman…

Cited by 72SourceScholar
2021

Real-Time Multi-View 3D Human Pose Estimation using Semantic Feedback to Smart Edge Sensors

RSS 2021poster

We present a novel method for estimation of 3D human poses from a multi-camera setup; employing distributed smart edge sensors coupled with a backend through a semantic feedback loop. 2D joint detection for each camera view is performed locally on a dedicated embedded inference processor. Only the…

2021

Real-time Multi-Adaptive-Resolution-Surfel 6D LiDAR Odometry using Continuous-time Trajectory Optimization

IROS 2021poster

Simultaneous Localization and Mapping (SLAM) is an essential capability for autonomous robots, but due to high data rates of 3D LiDARs real-time SLAM is challenging. We propose a real-time method for 6D LiDAR odometry. Our approach combines a continuous-time B-Spline trajectory representation with a…

Cited by 47SourcecodeScholar
2021

Search-based Planning of Dynamic MAV Trajectories Using Local Multiresolution State Lattices

ICRA 2021poster

Search-based methods that use motion primitives can incorporate the system’s dynamics into the planning and thus generate dynamically feasible MAV trajectories that are globally optimal. However, searching high-dimensional state lattices is computationally expensive. Local multiresolution is a commo…

Cited by 9SourceScholar
2020

Beyond Photometric Consistency: Gradient-based Dissimilarity for Improving Visual Odometry and Stereo Matching

ICRA 2020poster

Pose estimation and map building are central ingredients of autonomous robots and typically rely on the registration of sensor data. In this paper, we investigate a new metric for registering images that builds upon on the idea of the photometric error. Our approach combines a gradient orientation-b…

Cited by 6SourceScholar
2020

LatticeNet: Fast Point Cloud Segmentation Using Permutohedral Lattices

RSS 2020poster

Deep convolutional neural networks (CNNs) have shown outstanding performance in the task of semantically segmenting images. Applying the same methods on 3D data still poses challenges due to the heavy memory requirements and the lack of structured data. Here, we propose LatticeNet, a novel approach…

2019

A VR System for Immersive Teleoperation and Live Exploration with a Mobile Robot

IROS 2019poster

Applications like disaster management and industrial inspection often require experts to enter contaminated places. To circumvent the need for physical presence, it is desirable to generate a fully immersive individual live teleoperation experience. However, standard video-based approaches suffer fr…

Cited by 135SourceScholar
2019

Interpretable and Fine-Grained Visual Explanations for Convolutional Neural Networks

CVPR 2019poster

To verify and validate networks, it is essential to gain insight into their decisions, limitations as well as possible shortcomings of training data. In this work, we propose a post-hoc, optimization based visual explanation method, which highlights the evidence in the input image for a specific pre…

Cited by 186PDFScholar
2019

Search-based 3D Planning and Trajectory Optimization for Safe Micro Aerial Vehicle Flight Under Sensor Visibility Constraints

ICRA 2019poster

Safe navigation of Micro Aerial Vehicles (MAVs) requires not only obstacle-free flight paths according to a static environment map, but also the perception of and reaction to previously unknown and dynamic objects. This implies that the onboard sensors cover the current flight direction. Due to the…

Cited by 19SourceScholar
2019

SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences

ICCV 2019oral

Semantic scene understanding is important for various applications. In particular, self-driving cars need a fine-grained understanding of the surfaces and objects in their vicinity. Light detection and ranging (LiDAR) provides precise geometric information about the environment and is thus a part of…

Cited by 2413PDFcodeScholar
2019

Towards Learning Abstract Representations for Locomotion Planning in High-dimensional State Spaces

ICRA 2019poster

Ground robots which are able to navigate a variety of terrains are needed in many domains. One of the key aspects is the capability to adapt to the ground structure, which can be realized through movable body parts coming along with additional degrees of freedom (DoF). However, planning respective l…

Cited by 16SourceScholar
2018

Fast Autonomous Flight in Warehouses for Inventory Applications

RA-L 2018

The past years have shown a remarkable growth in use-cases for micro aerial vehicles (MAVs). Conceivable indoor applications require highly robust environment perception, fast reaction to changing situations, and stable navigation, but reliable sources of absolute positioning such as global navigati

Cited by 93SourceScholar
2018

Fast Object Learning and Dual-arm Coordination for Cluttered Stowing, Picking, and Packing

ICRA 2018poster

Robotic picking from cluttered bins is a demanding task, for which Amazon Robotics holds challenges. The 2017 Amazon Robotics Challenge (ARC) required stowing items into a storage system, picking specific items, and packing them into boxes. In this paper, we describe the entry of team NimbRo Picking…

Cited by 107SourceScholar
2018

Keyframe-Based Photometric Online Calibration and Color Correction

IROS 2018poster

Finding the parameters of a vignetting function for a camera currently involves the acquisition of several images in a given scene under very controlled lighting conditions, a cumbersome and error-prone task where the end result can only be confirmed visually. Many computer vision algorithms assume…

Cited by 3SourceScholar
2018

Robust 6D Object Pose Estimation in Cluttered Scenes Using Semantic Segmentation and Pose Regression Networks

IROS 2018poster

Object pose estimation is a crucial prerequisite for robots to perform autonomous manipulation in clutter. Real-world bin-picking settings such as warehouses present additional challenges, e.g., new objects are added constantly. Most of the existing object pose estimation methods assume that 3D mode…

Cited by 23SourceScholar
2018

Supervised Autonomous Locomotion and Manipulation for Disaster Response with a Centaur-Like Robot

IROS 2018poster

Mobile manipulation tasks are one of the key challenges in the field of search and rescue (SAR) robotics requiring robots with flexible locomotion and manipulation abilities. Since the tasks are mostly unknown in advance, the robot has to adapt to a wide variety of terrains and workspaces during a m…

Cited by 82SourceScholar
2018

Transferring Grasping Skills to Novel Instances by Latent Space Non-Rigid Registration

ICRA 2018poster

Robots acting in open environments need to be able to handle novel objects. Based on the observation that objects within a category are often similar in their shapes and usage, we propose an approach for transferring grasping skills from known instances to novel instances of an object category. Corr…

Cited by 39SourceScholar
2017

Combining Semantic and Geometric Features for Object Class Segmentation of Indoor Scenes

RA-L 2017

Scene understanding is a necessary prerequisite for robots acting autonomously in complex environments. Low-cost RGB-D cameras such as Microsoft Kinect enabled new methods for analyzing indoor scenes and are now ubiquitously used in indoor robotics. We investigate strategies for efficient pixelwise

Cited by 54SourceScholar
2017

NimbRo picking: Versatile part handling for warehouse automation

ICRA 2017poster

Part handling in warehouse automation is challenging if a large variety of items must be accommodated and items are stored in unordered piles. To foster research in this domain, Amazon holds picking challenges. We present our system which achieved second and third place in the Amazon Picking Challen…

Cited by 108SourceScholar
2017

Online depth calibration for RGB-D cameras using visual SLAM

IROS 2017poster

Modern consumer RGB-D cameras are affordable and provide dense depth estimates at high frame rates. Hence, they are popular for building dense environment representations. Yet, the sensors often do not provide accurate depth estimates since the factory calibration exhibits a static deformation. We p…

Cited by 9SourceScholar
2016

Efficient multi-camera visual-inertial SLAM for micro aerial vehicles

IROS 2016poster

Visual SLAM is an area of vivid research and bears countless applications for moving robots. In particular, micro aerial vehicles benefit from visual sensors due to their low weight. Their motion is, however, often faster and more complex than that of ground-based robots which is why systems with mu…

Cited by 49SourceScholar
2016

Hybrid driving-stepping locomotion with the wheeled-legged robot Momaro

ICRA 2016

Locomotion in uneven terrain is important for a wide range of robotic applications, including Search&Rescue operations. Our mobile manipulation robot Momaro features a unique locomotion design consisting of four legs ending in pairs of steerable wheels, allowing the robot to omnidirectionally drive

Cited by 61SourceScholar
2016

Local multiresolution trajectory optimization for micro aerial vehicles employing continuous curvature transitions

IROS 2016poster

Complex indoor and outdoor missions for autonomous micro aerial vehicles (MAV) require fast generation of collision-free paths in 3D space. Often not all obstacles in an environment are known prior to the mission execution. Consequently, the ability for replanning during a flight is key for success.…

Cited by 13SourceScholar
2015

RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features

ICRA 2015poster

Object recognition and pose estimation from RGB-D images are important tasks for manipulation robots which can be learned from examples. Creating and annotating datasets for learning is expensive, however. We address this problem with transfer learning from deep convolutional neural networks (CNN) t…

Cited by 437SourceScholar
2015

Real-time object detection, localization and verification for fast robotic depalletizing

IROS 2015poster

Depalletizing is a challenging task for manipulation robots. Key to successful application are not only robustness of the approach, but also achievable cycle times in order to keep up with the rest of the process. In this paper, we propose a system for depalletizing and a complete pipeline for detec…

Cited by 60SourceScholar