← Search

Stefano Rosa

7 accepted papers

2025

Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions

ICCV 2025poster

We present a self-supervised method to improve an agent's abilities in describing arbitrary objects while actively exploring a generic environment. This is a challenging problem, as current models struggle to obtain coherent image captions due to different camera viewpoints and clutter. We propose a…

2024

Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation

IROS 2024

Vision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks. Agents are tasked to navigate towards a target goal by executing a set of low-level actions, following a series of natural language instructions. All VLN-CE methods in t

Cited by 13SourceScholar
2020

DeepTIO: A Deep Thermal-Inertial Odometry With Visual Hallucination

RA-L 2020

Visual odometry shows excellent performance in a wide range of environments. However, in visually-denied scenarios (e.g. heavy smoke or darkness), pose estimates degrade or even fail. Thermal cameras are commonly used for perception and inspection when the environment has low visibility. However, th

Cited by 72SourceScholar
2020

RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds

CVPR 2020oral

We study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we i…

Cited by 2143PDFcodeScholar
2019

Selective Sensor Fusion for Neural Visual-Inertial Odometry

CVPR 2019poster

Deep learning approaches for Visual-Inertial Odometry (VIO) have proven successful, but they rarely focus on incorporating robust fusion strategies for dealing with imperfect input sensory data. We propose a novel end-to-end selective sensor fusion framework for monocular VIO, which fuses monocular…

Cited by 192PDFcodeScholar
2018

DEFO-NET: Learning Body Deformation Using Generative Adversarial Networks

ICRA 2018poster

Modelling the physical properties of everyday objects is a fundamental prerequisite for autonomous robots. We present a novel generative adversarial network (DEFO-NET), able to predict body deformations under external forces from a single RGB-D image. The network is based on an invertible conditiona…

Cited by 10SourceScholar
2018

Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning

ICRA 2018poster

Deep Reinforcement Learning (DRL) has been applied successfully to many robotic applications. However, the large number of trials needed for training is a key issue. Most of existing techniques developed to improve training efficiency (e.g. imitation) target on general tasks rather than being tailor…

Cited by 105SourcecodeScholar