← Search

Stefan Leutenegger

66 accepted papers

2026

Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation

ICRA 2026poster

We present Actron3D, a framework that enables robots to acquire transferable 6-DoF manipulation skills from monocular, uncalibrated, RGB-only human demonstration videos. Our key idea is to represent manipulation knowledge within a video as a continuous neural function over object space. At the core …

2026

FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment

ICRA 2026poster

Geometrically accurate and semantically expressive map representations have proven invaluable for robot deployment and task planning in unknown environments. Nevertheless, real-time, open-vocabulary semantic understanding of large-scale unknown environments still presents open challenges, mainly due…

2026

HumanFlow - Diffusion-Driven MAV Navigation Among Humans via Tightly-Coupled Motion Tracking, Forecasting, and Control

RSS 2026poster

Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches, however, often fail to predict plausible and accurate human motion estimates that are consistent with the surrounding scene, especially in the prese…

Cited by 0SourceScholar
2026

OKVIS2-X: Open Keyframe-Based Visual-Inertial SLAM Configurable with Dense Depth or LiDAR, and GNSS

ICRA 2026poster

To empower mobile robots with usable maps as well as highest state estimation accuracy and robustness, we present OKVIS2-X: a state-of-the-art multi-sensor Simultaneous Localization and Mapping (SLAM) system building dense volumetric occupancy maps, while scalable to large environments and operating…

2025

Digiforests: a Longitudinal Lidar Dataset for Forestry Robotics

ICRA 2025

Forests are vital to our ecosystems, acting as carbon sinks, climate stabilizers, biodiversity centers, and wood sources. Due to their scale, monitoring and managing forests takes a lot of work. Forestry robotics offers the potential for enabling efficient and sustainable foresting practices through

Cited by 12SourceScholar
2025

Efficient Submap-based Autonomous MAV Exploration using Visual-Inertial SLAM Configurable for LiDARs or Depth Cameras

ICRA 2025

Autonomous exploration of unknown space is an essential component for the deployment of mobile robots in the real world. Safe navigation is crucial for all robotics applications and requires accurate and consistent maps of the robot's surroundings. To achieve full autonomy and allow deployment in a

Cited by 6SourceScholar
2025

FrontierNet: Learning Visual Cues to Explore

RA-L 2025

Exploration of unknown environments is crucial for autonomous robots; it allows them to actively reason and decide on what new data to acquire for different tasks, such as mapping, object discovery, and environmental assessment. Existing solutions, such as frontier-based exploration approaches, rely

Cited by 12SourcecodeScholar
2025

REGRACE: A Robust and Efficient Graph-based Re-localization Algorithm using Consistency Evaluation

IROS 2025

Loop closures are essential for correcting odometry drift and creating consistent maps, especially in the context of large-scale navigation. Current methods using dense point clouds for accurate place recognition do not scale well due to computationally expensive scan-to-scan comparisons. Alternativ

Cited by 1SourceScholar
2025

Scalable Outdoors Autonomous Drone Flight with Visual-Inertial SLAM and Dense Submaps Built without LiDAR

IROS 2025

Autonomous navigation is needed for several robotics applications. In this paper we present an autonomous Micro Aerial Vehicle (MAV) system which purely relies on cost-effective and light-weight passive visual and inertial sensors to perform large-scale autonomous navigation in outdoor, unstructured

Cited by 3SourcecodeScholar
2025

SuperEvent: Cross-Modal Learning of Event-based Keypoint Detection for SLAM

ICCV 2025poster

Event-based keypoint detection and matching holds significant potential, enabling the integration of event sensors into highly optimized Visual SLAM systems developed for frame cameras over decades of research. Unfortunately, existing approaches struggle with the motion-dependent appearance of keypo…

2025

Uncertainty-Aware Visual-Inertial SLAM with Volumetric Occupancy Mapping

ICRA 2025

We propose visual-inertial simultaneous localization and mapping that tightly couples sparse reprojection errors, inertial measurement unit pre-integrals, and relative pose factors with dense volumetric occupancy mapping. Hereby depth predictions from a deep neural network are fused in a fully proba

Cited by 5SourceScholar
2025

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

CVPR 2025poster

Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundamentally does not scale well. We argue that learning from in-the-wild human videos…

Cited by 0SourcePDFScholar
2024

Control-Barrier-Aided Teleoperation with Visual-Inertial SLAM for Safe MAV Navigation in Complex Environments

ICRA 2024poster

In this paper, we consider a Micro Aerial Vehicle (MAV) system teleoperated by a non-expert and introduce a perceptive safety filter that leverages Control Barrier Functions (CBFs) in conjunction with Visual-Inertial Simultaneous Localization and Mapping (VI-SLAM) and dense 3D occupancy mapping to g…

Cited by 3SourceScholar
2024

Dynamic LiDAR Re-simulation using Compositional Neural Fields

CVPR 2024highlight

We introduce DyNFL a novel neural field-based approach for high-fidelity re-simulation of LiDAR scans in dynamic driving scenes. DyNFL processes LiDAR measurements from dynamic environments accompanied by bounding boxes of moving objects to construct an editable neural field. This field comprising s…

2024

FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object

ICRA 2024poster

We present FuncGrasp, a framework that can infer dense yet reliable grasp configurations for unseen objects using one annotated object and single-view RGB-D observation via categorical priors. Unlike previous works that only transfer a set of grasp poses, FuncGrasp aims to transfer infinite configur…

Cited by 4SourceScholar
2024

IN-Sight: Interactive Navigation through Sight

IROS 2024poster

Current visual navigation systems often treat the environment as static, lacking the ability to adaptively interact with obstacles. This limitation leads to navigation failure when encountering unavoidable obstructions. In response, we introduce IN-Sight, a novel approach to self-supervised path pla…

Cited by 2SourceScholar
2024

Online Tree Reconstruction and Forest Inventory on a Mobile Robotic System

IROS 2024poster

Terrestrial laser scanning (TLS) is the standard technique used to create accurate point clouds for digital forest inventories. However, the measurement process is demanding, requiring up to two days per hectare for data collection, significant data storage, as well as resource-heavy post-processing…

Cited by 8SourceScholar
2024

Tightly-Coupled LiDAR-Visual-Inertial SLAM and Large-Scale Volumetric Occupancy Mapping

ICRA 2024poster

Autonomous navigation is one of the key requirements for every potential application of mobile robots in the real-world. Besides high-accuracy state estimation, a suitable and globally consistent representation of the 3D environment is indispensable. We present a fully tightly-coupled LiDAR-Visual-I…

Cited by 8SourceScholar
2023

Accurate and Interactive Visual-Inertial Sensor Calibration with Next-Best-View and Next-Best-Trajectory Suggestion

IROS 2023poster

Visual-Inertial (VI) sensors are popular in robotics, self-driving vehicles, and augmented and virtual reality applications. In order to use them for any computer vision or state-estimation task, a good calibration is essential. However, collecting informative calibration data in order to render the…

Cited by 3SourcecodeScholar
2023

Anthropomorphic Grasping With Neural Object Shape Completion

RA-L 2023

The progressive prevalence of robots in human-suited environments has given rise to a myriad of object manipulation techniques, in which dexterity plays a paramount role. It is well-established that humans exhibit extraordinary dexterity when handling objects. Such dexterity seems to derive from a r

Cited by 12SourceScholar
2023

BodySLAM++: Fast and Tightly-Coupled Visual-Inertial Camera and Human Motion Tracking

IROS 2023poster

Robust, fast, and accurate human state - 6D pose and posture - estimation remains a challenging problem. For real-world applications, the ability to estimate the human state in realtime is highly desirable. In this paper, we present BodySLAM++, a fast, efficient, and accurate human and camera state…

Cited by 12SourceScholar
2023

Finding Things in the Unknown: Semantic Object-Centric Exploration with an MAV

ICRA 2023poster

Exploration of unknown space with an autonomous mobile robot is a well-studied problem. In this work we broaden the scope of exploration, moving beyond the pure geometric goal of uncovering as much free space as possible. We believe that for many practical applications, exploration should be context…

Cited by 23SourceScholar
2023

GloPro: Globally-Consistent Uncertainty-Aware 3D Human Pose Estimation & Tracking in the Wild

IROS 2023poster

An accurate and uncertainty-aware 3D human body pose estimation is key to enabling truly safe but efficient human-robot interactions. Current uncertainty-aware methods in 3D human pose estimation are limited to predicting the uncertainty of the body posture, while effectively neglecting the body sha…

Cited by 3SourceScholar
2023

Incremental Dense Reconstruction From Monocular Video With Guided Sparse Feature Volume Fusion

RA-L 2023

Incrementally recovering 3D dense structures from monocular videos is of paramount importance since it enables various robotics and AR applications. Feature volumes have recently been shown to enable efficient and accurate incremental dense reconstruction without the need to first estimate depth, bu

Cited by 12SourceScholar
2023

Orientation-Aware Hierarchical, Adaptive-Resolution A* Algorithm for UAV Trajectory Planning

RA-L 2023

Successful path planning for Unmanned Aerial Vehicles (UAVs) in challenging environments with narrow openings, such as disaster areas, requires attitude to be considered. State-of-the-art methods incorporate attitude only in the refinement stage. We introduce a first-of-a-kind global minimum cost pa

Cited by 14SourceScholar
2022

"BodySLAM: Joint Camera Localisation, Mapping, and Human Motion Tracking"

ECCV 2022poster

"Estimating human motion from video is an active research area due to its many potential applications. Most state-of-the-art methods predict human shape and posture estimates for individual images and do not leverage the temporal information available in video. Many ""in the wild"" sequences of huma…

2022

Learning to Complete Object Shapes for Object-level Mapping in Dynamic Scenes

IROS 2022poster

In this paper, we propose a novel object-level mapping system that can simultaneously segment, track, and reconstruct objects in dynamic scenes. It can further predict and complete their full geometries by conditioning on reconstructions from depth inputs and a category-level shape prior with the ai…

Cited by 13SourceScholar
2022

Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose Estimation

CVPR 2022poster

We propose a keypoint-based object-level SLAM framework that can provide globally consistent 6DoF pose estimates for symmetric and asymmetric objects alike. To the best of our knowledge, our system is among the first to utilize the camera pose information from SLAM to provide prior knowledge for tra…

Cited by 51PDFcodeScholar
2022

Visual-Inertial Multi-Instance Dynamic SLAM with Object-level Relocalisation

IROS 2022poster

In this paper, we present a tightly-coupled visual-inertial object-level multi-instance dynamic SLAM system. Even in extremely dynamic scenes, it can robustly optimise for the camera pose, velocity, IMU biases and build a dense 3D reconstruction object-level map of the environment. Our system can ro…

Cited by 16SourcecodeScholar
2022

Visual-Inertial SLAM with Tightly-Coupled Dropout-Tolerant GPS Fusion

IROS 2022poster

Robotic applications are continuously striving towards higher levels of autonomy. To achieve that goal, a highly robust and accurate state estimation is indispensable. Combining visual and inertial sensor modalities has proven to yield accurate and locally consistent results in short-term applicatio…

Cited by 20SourceScholar
2021

Elastic and Efficient LiDAR Reconstruction for Large-Scale Exploration Tasks

ICRA 2021poster

We present an efficient, elastic 3D LiDAR reconstruction framework which can reconstruct up to maximum Li-DAR ranges (60 m) at multiple frames per second, thus enabling robot exploration in large-scale environments. Our approach only requires a CPU. We focus on three main challenges of large-scale r…

Cited by 25SourceScholar
2021

In-Place Scene Labelling and Understanding With Implicit Scene Representation

ICCV 2021poster

Semantic labelling is highly correlated with geometry and radiance reconstruction, as scene entities with similar shape and appearance are more likely to come from similar classes. Recent implicit neural reconstruction techniques are appealing as they do not require prior training data, but the same…

Cited by 527PDFScholar
2021

Multi-Resolution 3D Mapping With Explicit Free Space Representation for Fast and Accurate Mobile Robot Motion Planning

RA-L 2021

With the aim of bridging the gap between high quality reconstruction and robot motion planning, we propose an efficient system that leverages the concept of adaptive-resolution volumetric mapping, which naturally integrates with the hierarchical decomposition of space in an octree data structure. In

Cited by 61SourceScholar
2021

SIMstack: A Generative Shape and Instance Model for Unordered Object Stacks

ICCV 2021poster

By estimating 3D shape and instances from a single view, we can capture information about the environment quickly, without the need for comprehensive scanning and multi-view fusion. Solving this task for composite scenes (such as object stacks) is challenging: occluded areas are not only ambiguous i…

Cited by 9PDFScholar
2021

Volumetric Occupancy Mapping With Probabilistic Depth Completion for Robotic Navigation

RA-L 2021

In robotic applications, a key requirement for safe and efficient motion planning is the ability to map obstacle-free space in unknown, cluttered 3D environments. However, commodity-grade RGB-D cameras commonly used for sensing fail to register valid depth values on shiny, glossy, bright, or distant

Cited by 27SourceScholar
2020

Aerial Manipulation Using Hybrid Force and Position NMPC Applied to Aerial Writing

RSS 2020poster

Aerial manipulation aims at combining the maneuverability of aerial vehicles with the manipulation capabilities of robotic arms. This, however, comes at the cost of the additional control complexity due to the coupling of the dynamics of the two systems. In this paper we present a Nonlinear Model Pr…

Cited by 70SourcePDFScholar
2020

Comparing View-Based and Map-Based Semantic Labelling in Real-Time SLAM

ICRA 2020poster

Generally capable Spatial AI systems must build persistent scene representations where geometric models are combined with meaningful semantic labels. The many approaches to labelling scenes can be divided into two clear groups: view-based which estimate labels from the input view-wise data and then…

Cited by 6SourceScholar
2020

Fast Frontier-based Information-driven Autonomous Exploration with an MAV

ICRA 2020poster

Exploration and collision-free navigation through an unknown environment is a fundamental task for autonomous robots. In this paper, a novel exploration strategy for Micro Aerial Vehicles (MAVs) is presented. The goal of the exploration strategy is the reduction of map entropy regarding occupancy pr…

Cited by 143SourceScholar
2020

Nonlinear MPC with Motor Failure Identification and Recovery for Safe and Aggressive Multicopter Flight

ICRA 2020poster

Safe and precise reference tracking is a crucial characteristic of Micro Aerial Vehicles (MAVs) that have to operate under the influence of external disturbances in cluttered environments. In this paper, we present a Nonlinear Model Predictive Control (NMPC) that exploits the fully physics based non…

Cited by 25SourceScholar
2020

Towards the Probabilistic Fusion of Learned Priors into Standard Pipelines for 3D Reconstruction

ICRA 2020poster

The best way to combine the results of deep learning with standard 3D reconstruction pipelines remains an open problem. While systems that pass the output of traditional multi-view stereo approaches to a network for regularisation or refinement currently seem to get the best results, it may be prefe…

Cited by 3SourceScholar
2019

Characterizing Visual Localization and Mapping Datasets

ICRA 2019poster

Benchmarking mapping and motion estimation algorithms is established practice in robotics and computer vision. As the diversity of datasets increases, in terms of the trajectories, models, and scenes, it becomes a challenge to select datasets for a given benchmarking purpose. Inspired by the Wassers…

Cited by 30SourceScholar
2019

DeepFusion: Real-Time Dense 3D Reconstruction for Monocular SLAM using Single-View Depth and Gradient Predictions

ICRA 2019poster

While the keypoint-based maps created by sparse monocular Simultaneous Localisation and Mapping (SLAM) systems are useful for camera tracking, dense 3D reconstructions may be desired for many robotic tasks. Solutions involving depth cameras are limited in range and to indoor spaces, and dense recons…

Cited by 70SourceScholar
2019

KO-Fusion: Dense Visual SLAM with Tightly-Coupled Kinematic and Odometric Tracking

ICRA 2019poster

Dense visual SLAM methods are able to estimate the 3D structure of an environment and locate the observer within them. They estimate the motion of a camera by matching visual information between consecutive frames, and are thus prone to failure under extreme motion conditions or when observing textu…

Cited by 19SourceScholar
2019

Learning Meshes for Dense Visual SLAM

ICCV 2019poster

Estimating motion and surrounding geometry of a moving camera remains a challenging inference problem. From an information theoretic point of view, estimates should get better as more information is included, such as is done in dense SLAM, but this is strongly dependent on the validity of the underl…

Cited by 28PDFScholar
2019

MID-Fusion: Octree-based Object-Level Multi-Instance Dynamic SLAM

ICRA 2019poster

We propose a new multi-instance dynamic RGB-D SLAM system using an object-level octree-based volumetric representation. It can provide robust camera tracking in dynamic environments and at the same time, continuously estimate geometric, semantic, and motion properties for arbitrary objects in the sc…

Cited by 240SourcecodeScholar
2019

SceneCode: Monocular Dense Semantic Reconstruction Using Learned Encoded Scene Representations

CVPR 2019poster

Systems which incrementally create 3D semantic maps from image sequences must store and update representations of both geometry and semantic entities. However, while there has been much work on the correct formulation for geometrical estimation, state-of-the-art systems usually rely on simple semant…

Cited by 95PDFScholar
2018

CodeSLAM — Learning a Compact, Optimisable Representation for Dense Visual SLAM

CVPR 2018poster

The representation of geometry in real-time 3D perception systems continues to be a critical research issue. Dense maps capture complete surface shape and can be augmented with semantic labels, but their high dimensionality makes them computationally costly to store and process, and unsuitable for r…

Cited by 463SourcePDFScholar
2018

Efficient Octree-Based Volumetric SLAM Supporting Signed-Distance and Occupancy Mapping

RA-L 2018

We present a dense volumetric simultaneous localisation and mapping (SLAM) framework that uses an octree representation for efficient fusion and rendering of either a truncated signed distance field (TSDF) or an occupancy map. The primary aim of this letter is to use one single representation of the

Cited by 127SourceScholar
2018

Learning to Solve Nonlinear Least Squares for Monocular Stereo

ECCV 2018poster

Sum-of-squares objective functions are very popular in computer vision algorithms. However, these objective functions are not always easy to optimize. The underlying assumptions made by solvers are often not satisfied and many problems are inherently ill-posed. In this paper, we propose a neural non…

Cited by 100SourcePDFScholar
2017

Monocular visual odometry: Sparse joint optimisation or dense alternation?

ICRA 2017poster

Real-time monocular SLAM is increasingly mature and entering commercial products. However, there is a divide between two techniques providing similar performance. Despite the rise of ‘dense’ and ‘semi-dense’ methods which use large proportions of the pixels in a video stream to estimate motion and s…

Cited by 22SourceScholar
2017

SceneNet RGB-D: Can 5M Synthetic Images Beat Generic ImageNet Pre-Training on Indoor Segmentation?

ICCV 2017poster

We introduce SceneNet RGB-D, a dataset providing pixel-perfect ground truth for scene understanding problems such as semantic segmentation, instance segmentation, and object detection. It also provides perfect camera poses and depth data, allowing investigation into geometric computer vision problem…

Cited by 362PDFcodeScholar
2017

SemanticFusion: Dense 3D semantic mapping with convolutional neural networks

ICRA 2017poster

Ever more robust, accurate and detailed mapping using visual sensing has proven to be an enabling factor for mobile robots across a wide variety of applications. For the next level of robot intelligence and intuitive user interaction, maps need to extend beyond geometry and appearance - they need to…

Cited by 830SourceScholar
2016

Deep learning a grasp function for grasping under gripper pose uncertainty

IROS 2016poster

This paper presents a new method for parallel-jaw grasping of isolated objects from depth images, under large gripper pose uncertainty. Whilst most approaches aim to predict the single best grasp pose from an image, our method first predicts a score for every possible grasp pose, which we denote the…

Cited by 307SourceScholar
2016

Pairwise Decomposition of Image Sequences for Active Multi-View Recognition

CVPR 2016oral

A multi-view image sequence provides a much richer capacity for object recognition than from a single image. However, most existing solutions to multi-view recognition typically adopt hand-crafted, model-based geometric methods, which do not readily embrace recent trends in deep learning. We propose…

Cited by 308PDFScholar
2016

Simultaneous Optical Flow and Intensity Estimation From an Event Camera

CVPR 2016oral

Event cameras are bio-inspired vision sensors which mimic retinas to measure per-pixel intensity change rather than outputting an actual intensity image. This proposed paradigm shift away from traditional frame cameras offers significant potential advantages: namely avoiding high data rates, dynamic…

Cited by 377PDFScholar
2015

A solar-powered hand-launchable UAV for low-altitude multi-day continuous flight

ICRA 2015poster

This paper presents the conceptual design, detailed development and flight testing of AtlantikSolar, a 5.6m-wingspan solar-powered Low-Altitude Long-Endurance (LALE) Unmanned Aerial Vehicle (UAV) designed and built at ETH Zurich. The UAV is required to provide perpetual endurance at a geographic lat…

Cited by 118SourceScholar
2015

ElasticFusion: Dense SLAM Without A Pose Graph

RSS 2015poster

We present a novel approach to real-time dense visual SLAM. Our system is capable of capturing comprehensive dense globally consistent surfel-based maps of room scale environments explored using an RGB-D camera in an incremental online fashion, without pose graph optimisation or any post-processing…

Cited by 1054SourcePDFScholar