← Search

Cesar Cadena

69 accepted papers

2026

BIEVR-LIO: Robust LiDAR-Inertial Odometry through Bump-Image-Enhanced Voxel Maps

RSS 2026poster

Reliable odometry is essential for mobile robots as they increasingly enter more challenging environments, which often contain little information to constrain point cloud registration, resulting in degraded LiDAR–Inertial Odometry (LIO) accuracy or even divergence. To address this, we present BIEVR-…

Cited by 0SourceScholar
2026

Informed, Constrained, Aligned: A Field Analysis on Degeneracy-Aware Point Cloud Registration in the Wild (I)

ICRA 2026poster

The iterative closest point registration algorithm has been a preferred method for light detection and ranging LiDAR-based robot localization for nearly a decade. However, even in modern simultaneous localization and mapping (SLAM) solutions, ICP can degrade and become unreliable in geometrically il…

Cited by 0Scholar
2026

Keypoint Semantic Integration for Improved Feature Matching in Outdoor Agricultural Environments

ICRA 2026poster

Robust robot navigation in outdoor environments requires accurate perception systems capable of handling visual challenges such as repetitive structures and changing appearances. Visual feature matching is crucial to vision-based pipelines but remains particularly challenging in natural outdoor sett…

2026

NaviTrace: Evaluating Embodied Navigation of Vision-Language Models

ICRA 2026poster

Vision–language models demonstrate unprecedented performance and generalization across a wide range of tasks and scenarios. Integrating these foundation models into robotic navigation systems opens pathways toward building general-purpose robots. Yet, evaluating these models’ navigation capabilities…

2026

Sight Over Site: Perception-Aware Reinforcement Learning for Efficient Robotic Inspection

ICRA 2026poster

Autonomous inspection is a central problem in robotics, with applications ranging from industrial monitoring to search-and-rescue. Traditionally, inspection has often been reduced to navigation tasks, where the objective is to reach a predefined location while avoiding obstacles. However, this formu…

2025

Boxi: Design Decisions in the Context of Algorithmic Performance for Robotics

RSS 2025poster

Achieving robust autonomy in mobile robots operating in complex, unstructured environments requires a multimodal sensor suite capable of capturing diverse and complementary information. However, designing such a sensor suite involves multiple critical design decisions, such as sensor selection, comp…

Cited by 1PDFScholar
2025

Efficient Hierarchical Any-Angle Path Planning on Multi-Resolution 3D Grids

RSS 2025poster

Hierarchical, multi-resolution volumetric mapping approaches are widely used to represent large and complex environments as they can efficiently capture their occupancy and connectivity information. Yet widely used path planning methods such as sampling and trajectory optimization do not exploit thi…

Cited by 0PDFScholar
2025

ForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density Images

CVPR 2025highlight

Place recognition is essential to maintain global consistency in large-scale localization systems. While research in urban environments has progressed significantly using LiDARs or cameras, applications in natural forest-like environments remain largely underexplored. Furthermore, forests present pa…

2025

FrontierNet: Learning Visual Cues to Explore

RA-L 2025

Exploration of unknown environments is crucial for autonomous robots; it allows them to actively reason and decide on what new data to acquire for different tasks, such as mapping, object discovery, and environmental assessment. Existing solutions, such as frontier-based exploration approaches, rely

Cited by 12SourcecodeScholar
2025

Learned Perceptive Forward Dynamics Model for Safe and Platform-aware Robotic Navigation

RSS 2025poster

Ensuring safe navigation in complex environments requires accurate real-time traversability assessment and understanding of environmental interactions relative to the robot’s capabilities. Traditional methods, which assume simplified dynamics, often require designing and tuning cost functions to saf…

Cited by 0PDFcodeScholar
2025

TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation

IROS 2025

We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in various photorealistic simulation environments includes multiple RGB stereo cameras for 360-degree coverage, along with de

Cited by 17SourceScholar
2024

COIN-LIO: Complementary Intensity-Augmented LiDAR Inertial Odometry

ICRA 2024poster

We present COIN-LIO, a LiDAR Inertial Odometry pipeline that tightly couples information from LiDAR intensity with geometry-based point cloud registration. The focus of our work is to improve the robustness of LiDAR-inertial odometry in geometrically degenerate scenarios, like tunnels or flat fields…

Cited by 23SourcecodeScholar
2024

Resilient Legged Local Navigation: Learning to Traverse with Compromised Perception End-to-End

ICRA 2024poster

Autonomous robots must navigate reliably in unknown environments even under compromised exteroceptive perception, or perception failures. Such failures often occur when harsh environments lead to degraded sensing, or when the perception algorithm misinterprets the scene due to limited generalization…

Cited by 16SourceScholar
2024

Tag Map: A Text-Based Map for Spatial Reasoning and Navigation with Large Language Models

CoRL 2024poster

Large Language Models (LLM) have emerged as a tool for robots to generate task plans using common sense reasoning. For the LLM to generate actionable plans, scene context must be provided, often through a map. Recent works have shifted from explicit maps with fixed semantic classes to implicit open…

Cited by 3SourceScholar
2024

Temporal- and Viewpoint-Invariant Registration for Under-Canopy Footage using Deep-Learning-based Bird’s-Eye View Prediction

IROS 2024poster

Conducting visual assessments under the canopy using mobile robots is an emerging task in smart farming and forestry. However, it is challenging to register images across different data-collection days, especially across seasons, due to the self-occluding geometry and temporal dynamics in forests an…

Cited by 1SourcecodeScholar
2024

VIRUS-NeRF - Vision, InfraRed and UltraSonic based Neural Radiance Fields

IROS 2024poster

Autonomous mobile robots are an increasingly integral part of modern factory and warehouse operations. Obstacle detection, avoidance and path planning are critical safety-relevant tasks, which are often solved using expensive LiDAR sensors and depth cameras. We propose to use cost-effective low-reso…

Cited by 2SourcecodeScholar
2023

3D VSG: Long-term Semantic Scene Change Prediction through 3D Variable Scene Graphs

ICRA 2023poster

Numerous applications require robots to operate in environments shared with other agents, such as humans or other robots. However, such shared scenes are typically subject to different kinds of long-term semantic scene changes. The ability to model and predict such changes is thus crucial for robot…

Cited by 27SourcecodeScholar
2023

Efficient volumetric mapping of multi-scale environments using wavelet-based compression

RSS 2023poster

Volumetric maps are widely used in robotics due to their desirable properties in applications such as path planning, exploration, and manipulation. Constant advances in mapping technologies are needed to keep up with the improvements in sensor technology, generating increasingly vast amounts of prec…

2023

Fast Traversability Estimation for Wild Visual Navigation

RSS 2023poster

Natural environments such as forests and grasslands are challenging for robotic navigation because of the false perception of rigid obstacles from high grass, twigs, or bushes. In this work, we propose Wild Visual Navigation (WVN), an online self-supervised learning system for traversability estimat…

Cited by 78SourcePDFScholar
2023

Local and Global Information in Obstacle Detection on Railway Tracks

IROS 2023poster

Reliable obstacle detection on railways could help prevent collisions that result in injuries and potentially damage or derail the train. Unfortunately, generic object detectors do not have enough classes to account for all possible scenarios, and datasets featuring objects on railways are challengi…

Cited by 11SourceScholar
2023

Obstacle avoidance using Raycasting and Riemannian Motion Policies at kHz rates for MAVs

ICRA 2023poster

This paper presents a novel method for using Riemannian Motion Policies on volumetric maps, shown in the example of obstacle avoidance for Micro Aerial Vehicles (MAVs), Today, most robotic obstacle avoidance algorithms rely on sampling or optimization-based planners with volumetric maps. However, th…

Cited by 16SourcecodeScholar
2023

Seeing Through the Grass: Semantic Pointcloud Filter for Support Surface Learning

RA-L 2023

Mobile ground robots require perceiving and understanding their surrounding support surface to move around autonomously and safely. The support surface is commonly estimated based on exteroceptive depth measurements, e.g., from LiDARs. However, the measured depth fails to align with the true support

Cited by 18SourceScholar
2023

SphNet: A Spherical Network for Semantic Pointcloud Segmentation

ICRA 2023poster

Semantic segmentation for robotic systems can enable a wide range of applications, from self-driving cars and augmented reality systems to domestic robots. We argue that a spherical representation is a natural one for egocentric pointclouds. Thus, in this work, we present a novel framework exploitin…

Cited by 2SourceScholar
2023

Unsupervised Continual Semantic Adaptation Through Neural Rendering

CVPR 2023poster

An increasing amount of applications rely on data-driven models that are deployed for perception tasks across a sequence of scenes. Due to the mismatch between training and deployment data, adapting the model on the new scenes is often crucial to obtain good performance. In this work, we study conti…

2023

maplab 2.0 - A Modular and Multi-Modal Mapping Framework

RA-L 2023

Integration of multiple sensor modalities and deep learning into Simultaneous Localization And Mapping (SLAM) systems are areas of significant interest in current research. Multi-modality is a stepping stone towards achieving robustness in challenging environments and interoperability of heterogeneo

Cited by 77SourcecodeScholar
2022

Collaborative Robot Mapping using Spectral Graph Analysis

ICRA 2022poster

In this paper, we deal with the problem of creating globally consistent pose graphs in a centralized multi-robot SLAM framework. For each robot to act autonomously, individual onboard pose estimates and maps are maintained, which are then communicated to a central server to build an optimized global…

Cited by 15SourceScholar
2022

Continual Adaptation of Semantic Segmentation Using Complementary 2D-3D Data Representations

RA-L 2022

Semantic segmentation networks are usually pre-trained once and not updated during deployment. As a consequence, misclassifications commonly occur if the distribution of the training data deviates from the one encountered during the robot's operation. We propose to mitigate this problem by adapting

Cited by 16SourceScholar
2022

Don't Share My Face: Privacy Preserving Inpainting for Visual Localization

IROS 2022poster

Visual localization is an important task for many robotic and augmented reality applications. As localizing within large scale maps can be memory and computationally de-manding, cloud-based localization services are appealing for developers. However, such services raise important privacy concerns fo…

Cited by 4SourceScholar
2022

Embodied Active Domain Adaptation for Semantic Segmentation via Informative Path Planning

RA-L 2022

This work presents an embodied agent that can adapt its semantic segmentation network to new indoor environments in a fully autonomous way. Because semantic segmentation networks fail to generalize well to unseen environments, the agent collects images of the new environment which are then used for

Cited by 23SourcecodeScholar
2022

Panoptic Multi-TSDFs: a Flexible Representation for Online Multi-resolution Volumetric Mapping and Long-term Dynamic Scene Consistency

ICRA 2022poster

For robotic interaction in environments shared with other agents, access to volumetric and semantic maps of the scene is crucial. However, such environments are inevitably subject to long-term changes, which the map needs to account for. We thus propose panoptic multi-TSDFs as a novel representation…

Cited by 74SourcecodeScholar
2022

See Yourself in Others: Attending Multiple Tasks for Own Failure Detection

ICRA 2022poster

Autonomous robots deal with unexpected scenarios in real environments. Given input images, various visual perception tasks can be performed, e.g., semantic segmentation, depth estimation and normal estimation. These different tasks provide rich information for the whole robotic perception system. Al…

Cited by 12SourcecodeScholar
2021

Mesh Manifold Based Riemannian Motion Planning for Omnidirectional Micro Aerial Vehicles

RA-L 2021

This letter presents a novel on-line path planning method that enables aerial robots to interact with surfaces. We present a solution to the problem of finding trajectories that drive a robot towards a surface and move along it. Triangular meshes are used as a surface map representation that is free

Cited by 14SourceScholar
2021

PHASER: A Robust and Correspondence-Free Global Pointcloud Registration

RA-L 2021

We propose PHASER, a correspondence-free global registration of sensor-centric pointclouds that is robust to noise, sparsity, and partial overlaps. Our method can seamlessly handle multimodal information, and does not rely on keypoint nor descriptor preprocessing modules. By exploiting properties of

Cited by 37SourcecodeScholar
2021

Pixel-Wise Anomaly Detection in Complex Driving Scenes

CVPR 2021poster

The inability of state-of-the-art semantic segmentation methods to detect anomaly instances hinders them from being deployed in safety-critical and complex applications, such as autonomous driving. Recent approaches have focused on either leveraging segmentation uncertainty to identify anomalous are…

Cited by 173PDFcodeScholar
2021

Self-Improving Semantic Perception for Indoor Localisation

CoRL 2021poster

We propose a novel robotic system that can improve its perception during deployment. Contrary to the established approach of learning semantics from large datasets and deploying fixed models, we propose a framework in which semantic models are continuously updated on the robot to adapt to the deploy…

Cited by 8SourcecodeScholar
2021

Spherical Multi-Modal Place Recognition for Heterogeneous Sensor Systems

ICRA 2021poster

In this paper, we propose a robust end-to-end multi-modal pipeline for place recognition where the sensor systems can differ from the map building to the query. Our approach operates directly on images and LiDAR scans without requiring any local feature extraction modules. By projecting the sensor d…

Cited by 23SourcecodeScholar
2020

Depth Based Semantic Scene Completion With Position Importance Aware Loss

RA-L 2020

Semantic scene completion (SSC) refers to the task of inferring the 3D semantic segmentation of a scene while simultaneously completing the 3D shapes. We propose PALNet, a novel hybrid network for SSC based on single depth. PALNet utilizes a two-stream network to extract both 2D and 3D features from

Cited by 71SourcecodeScholar
2020

Learning Camera Miscalibration Detection

ICRA 2020poster

Self-diagnosis and self-repair are some of the key challenges in deploying robotic platforms for long-term real-world applications. One of the issues that can occur to a robot is miscalibration of its sensors due to aging, environmental transients, or external disturbances. Precise calibration lies…

Cited by 20SourcecodeScholar
2020

Learning Densities in Feature Space for Reliable Segmentation of Indoor Scenes

RA-L 2020

Deep learning has enabled remarkable advances in scene understanding, particularly in semantic segmentation tasks. Yet, current state of the art approaches are limited to a closed set of classes, and fail when facing novel elements, also known as out of distribution (OoD) data. This is a problem as

Cited by 21SourceScholar
2020

Learning Trajectories for Visual-Inertial System Calibration via Model-based Heuristic Deep Reinforcement Learning

CoRL 2020

Visual-inertial systems rely on precise calibrations of both camera intrinsics and inter-sensor extrinsics, which typically require manually performing complex motions in front of a calibration target. In this work we present a novel approach to obtain favorable trajectories for visual-inertial syst

2020

MOZARD: Multi-Modal Localization for Autonomous Vehicles in Urban Outdoor Environments

IROS 2020poster

Visually poor scenarios are one of the main sources of failure in visual localization systems in outdoor environments. To address this challenge, we present MOZARD, a multi-modal localization system for urban outdoor environments using vision and LiDAR. By fusing key point based visual multi-session…

Cited by 2SourceScholar
2020

Voxgraph: Globally Consistent, Volumetric Mapping Using Signed Distance Function Submaps

RA-L 2020

Globally consistent dense maps are a key requirement for long-term robot navigation in complex environments. While previous works have addressed the challenges of dense mapping and global consistency, most require more computational resources than may be available on-board small robots. We propose a

Cited by 109SourcecodeScholar
2019

An Approach for Semantic Segmentation of Tree-like Vegetation

ICRA 2019poster

This paper presents a pipeline for semantic segmentation of trees into their components. Given a single RGB-D image of a tree, we employ a deep network to predict labels to classify each pixel of the tree into trunk, branches, twigs and leaves. Multiple convolutional neural network architectures to…

Cited by 18SourceScholar
2019

Empty Cities: Image Inpainting for a Dynamic-Object-Invariant Space

ICRA 2019poster

In this paper we present an end-to-end deep learning framework to turn images that show dynamic content, such as vehicles or pedestrians, into realistic static frames. This objective encounters two main challenges: detecting all the dynamic objects, and inpainting the static occluded background with…

Cited by 38SourcecodeScholar
2019

Experimental Comparison of Visual-Aided Odometry Methods for Rail Vehicles

RA-L 2019

Today, rail vehicle localization is based on infrastructure-side Balises (beacons) together with on-board odometry to determine whether a rail segment is occupied. Such a coarse locking leads to a sub-optimal usage of the rail networks. New railway standards propose the use of moving blocks centred

Cited by 41SourceScholar
2019

Flexible Trinocular: Non-rigid Multi-Camera-IMU Dense Reconstruction for UAV Navigation and Mapping

IROS 2019poster

In this paper, we propose a visual-inertial framework able to efficiently estimate the camera poses of a non-rigid trinocular baseline for long-range depth estimation on-board a fast moving aerial platform. The estimation of the time-varying baseline is based on relative inertial measurements, a pho…

Cited by 9SourceScholar
2019

From Coarse to Fine: Robust Hierarchical Localization at Large Scale

CVPR 2019poster

Robust and accurate visual localization is a fundamental capability for numerous applications, such as autonomous driving, mobile robotics, or augmented reality. It remains, however, a challenging task, particularly for large-scale environments and in presence of significant appearance changes. Stat…

Cited by 1086PDFcodeScholar
2019

OREOS: Oriented Recognition of 3D Point Clouds in Outdoor Scenarios

IROS 2019poster

We introduce a novel method for oriented place recognition with 3D LiDAR scans. A Convolutional Neural Network is trained to extract compact descriptors from single 3D LiDAR scans. These can be used both to retrieve near-by place candidates from a map, and to estimate the yaw discrepancy needed for…

Cited by 64SourceScholar
2019

Object Classification Based on Unsupervised Learned Multi-Modal Features For Overcoming Sensor Failures

ICRA 2019poster

For autonomous driving applications it is critical to know which type of road users and road side infrastructure are present to plan driving manoeuvres accordingly. Therefore autonomous cars are equipped with different sensor modalities to robustly perceive its environment. However, for classificati…

Cited by 4SourceScholar
2019

Volumetric Instance-Aware Semantic Mapping and 3D Object Discovery

RA-L 2019

To autonomously navigate and plan interactions in real-world environments, robots require the ability to robustly perceive and map complex, unstructured surrounding scenes. Besides building an internal representation of the observed scene geometry, the key insight toward a truly functional understan

Cited by 255SourcecodeScholar
2019

Where Should I Walk? Predicting Terrain Properties From Images Via Self-Supervised Learning

RA-L 2019

Legged robots have the potential to traverse diverse and rugged terrain. To find a safe and efficient navigation path and to carefully select individual footholds, it is useful to be able to predict properties of the terrain ahead of the robot. In this letter, we propose a method to collect data fro

Cited by 206SourceScholar
2018

A Data-driven Model for Interaction-Aware Pedestrian Motion Prediction in Object Cluttered Environments

ICRA 2018poster

This paper reports on a data-driven, interaction-aware motion prediction approach for pedestrians in environments cluttered with static obstacles. When navigating in such workspaces shared with humans, robots need accurate motion predictions of the surrounding pedestrians. Human navigation behavior…

Cited by 140SourceScholar
2018

Automatic Segmentation of Tree Structure From Point Cloud Data

RA-L 2018

Methods for capturing and modeling vegetation, such as trees or plants, typically distinguish between two components-branch skeleton and foliage. Current methods do not provide quantitatively accurate tree structure and foliage density needed for applications such as visualization, inspection, or to

Cited by 22SourceScholar
2018

Free LSD: Prior-Free Visual Landing Site Detection for Autonomous Planes

RA-L 2018

Full autonomy for fixed-wing unmanned aerial vehicles (UAVs) requires the capability to autonomously detect potential landing sites in unknown and unstructured terrain, allowing for self-governed mission completion or handling of emergency situations. In this letter, we propose a perception system a

Cited by 39SourceScholar
2018

Incremental-Segment-Based Localization in 3-D Point Clouds

RA-L 2018

Localization in 3-D point clouds is a highly challenging task due to the complexity associated with extracting information from 3-D data. This letter proposes an incremental approach addressing this problem efficiently. The presented method first accumulates the measurements in a dynamic voxel grid

Cited by 61SourcecodeScholar
2018

Leveraging Deep Visual Descriptors for Hierarchical Efficient Localization

CoRL 2018

Many robotics applications require precise pose estimates despite operating in large and changing environments. This can be addressed by visual localization, using a pre-computed 3D model of the surroundings. The pose estimation then amounts to finding correspondences between 2D keypoints in a query

2018

Reinforced Imitation: Sample Efficient Deep Reinforcement Learning for Mapless Navigation by Leveraging Prior Demonstrations

RA-L 2018

This letter presents a case study of a learning-based approach for target-driven mapless navigation. The underlying navigation model is an end-to-end neural network, which is trained using a combination of expert demonstrations, imitation learning (IL) and reinforcement learning (RL). While RL and I

Cited by 176SourcecodeScholar
2018

SegMap: 3D Segment Mapping using Data-Driven Descriptors

RSS 2018poster

When performing localization and mapping, working at the level of structure can be advantageous in terms of robustness to environmental changes and differences in illumination. This paper presents SegMap: a map representation solution to the localization and mapping problem based on the extraction o…

2017

An online multi-robot SLAM system for 3D LiDARs

IROS 2017poster

Using multiple cooperative robots is advantageous for time critical Search and Rescue (SaR) missions as they permit rapid exploration of the environment and provide higher redundancy than using a single robot. A considerable number of applications such as autonomous driving and disaster response cou…

Cited by 176SourceScholar
2017

From perception to decision: A data-driven approach to end-to-end motion planning for autonomous ground robots

ICRA 2017poster

Learning from demonstration for motion planning is an ongoing research topic. In this paper we present a model that is able to learn the complex mapping from raw 2D-laser range findings and a target position to the required steering commands for the robot. To our best knowledge, this work presents t…

Cited by 526SourceScholar
2017

SegMatch: Segment based place recognition in 3D point clouds

ICRA 2017poster

Place recognition in 3D data is a challenging task that has been commonly approached by adapting image-based solutions. Methods based on local features suffer from ambiguity and from robustness to environment changes while methods based on global features are viewpoint dependent. We propose SegMatch…

Cited by 418SourceScholar
2017

TSDF-based change detection for consistent long-term dense reconstruction and dynamic object discovery

ICRA 2017poster

Robots that are operating for extended periods of time need to be able to deal with changes in their environment and represent them adequately in their maps. In this paper, we present a novel 3D reconstruction algorithm based on an extended Truncated Signed Distance Function (TSDF) that enables to c…

Cited by 93SourceScholar
2016

Multi-modal Auto-Encoders as Joint Estimators for Robotics Scene Understanding

RSS 2016poster

We explore the capabilities of Auto-Encoders to fuse the information available from cameras and depth sensors, and to reconstruct missing data, for scene understanding tasks. In particular we consider three input modalities: RGB images; depth images; and semantic label information. We seek to genera…

Cited by 122SourcePDFScholar
2015

A fast, modular scene understanding system using context-aware object detection

ICRA 2015poster

We propose a semantic scene understanding system that is suitable for real robotic operations. The system solves different tasks (semantic segmentation and object detections) in an opportunistic and distributed fashion but still allows communication between modules to improve their respective perfor…

Cited by 37SourceScholar