← Search

Mirko Nava

10 accepted papers

2026

Self-Supervised Domain Adaptation for Visual 3D Pose Estimation of Nano-Drone Racing Gates by Enforcing Geometric Consistency

ICRA 2026poster

We consider the task of visually estimating the relative pose of a drone racing gate in front of a nano-quadrotor, using a convolutional neural network pre-trained on simulated data to regress the gate's pose. Due to the sim-to-real gap, the pre-trained model underperforms in the real world and must…

2025

Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States

CoRL 2025poster

We introduce a model for monocular RGB relative pose estimation of a ground robot that trains from scratch without pose labels nor prior knowledge about the robot's shape or appearance. At training time, we assume: (i) a robot fitted with multiple LEDs, whose states are independent and known at each…

Cited by 0SourceScholar
2024

Learning to Estimate the Pose of a Peer Robot in a Camera Image by Predicting the States of its LEDs

IROS 2024poster

We consider the problem of training a fully convolutional network to estimate the relative 6D pose of a robot given a camera image, when the robot is equipped with independent controllable LEDs placed in different parts of its body. The training data is composed by few (or zero) images labeled with…

Cited by 0SourceScholar
2024

Self-Supervised Learning of Visual Robot Localization Using LED State Prediction as a Pretext Task

RA-L 2024

We propose a novel self-supervised approach for learning to visually localize robots equipped with controllable LEDs. We rely on a few training samples labeled with position ground truth and many training samples in which only the LED state is known, whose collection is cheap. We show that using LED

Cited by 4SourcecodeScholar
2022

Learning Visual Localization of a Quadrotor Using Its Noise as Self-Supervision

RA-L 2022

We introduce an approach to train neural network models for visual object localization using a small training set, labeled with ground truth object positions and a large unlabeled one. We assume that the object to be localized emits sound, which is perceived by a microphone rigidly affixed to the ca

Cited by 14SourceScholar
2022

Visual Servoing with Geometrically Interpretable Neural Perception

IROS 2022poster

An increasing number of nonspecialist robotic users demand easy-to-use machines. In the context of visual servoing, the removal of explicit image processing is becoming a trend, allowing an easy application of this technique. This work presents a deep learning approach for solving the perception pro…

Cited by 6SourceScholar
2021

State-Consistency Loss for Learning Spatial Perception Tasks From Partial Labels

RA-L 2021

When learning models for real-world robot spatial perception tasks, one might have access only to partial labels: this occurs for example in semi-supervised scenarios (in which labels are not available for a subset of the training instances) or in some types of self-supervised robot learning (where

Cited by 7SourceScholar
2021

Uncertainty-Aware Self-Supervised Learning of Spatial Perception Tasks

RA-L 2021

We propose a general self-supervised learning approach for spatial perception tasks, such as estimating the pose of an object relative to the robot, from onboard sensor readings. The model is learned from training episodes, by relying on: A continuous state estimate, possibly inaccurate and affected

Cited by 17SourcecodeScholar
2020

Path Planning With Local Motion Estimations

RA-L 2020

We introduce a novel approach to long-range path planning that relies on a learned model to predict the outcome of local motions using possibly partial knowledge. The model is trained from a dataset of trajectories acquired in a self-supervised way. Sampling-based path planners use this component to

Cited by 46SourceScholar
2019

Learning Long-Range Perception Using Self-Supervision From Short-Range Sensors and Odometry

RA-L 2019

We introduce a general self-supervised approach to predict the future outputs of a short-range sensor (such as a proximity sensor) given the current outputs of a long-range sensor (such as a camera). We assume that the former is directly related to some piece of information to be perceived (such as

Cited by 29SourceScholar