← Search

Daniele De Martini

28 accepted papers

2026

3D Foundation Model-Based Loop Closing for Decentralized Collaborative SLAM

ICRA 2026poster

Decentralized Collaborative Simultaneous Localization And Mapping (C-SLAM) techniques often struggle to identify map overlaps due to significant viewpoint variations among robots. Motivated by recent advancements in 3D foundation models, which can register images despite large viewpoint differences,…

2026

Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference

ICLR 2026poster

Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in…

Cited by 0SourcecodeScholar
2026

Select2Plan: Training-Free ICL-Based Planning through VQA and Memory Retrieval

ICRA 2026poster

We introduce Select2Plan (S2P), a novel training-free framework for high-level robot planning that leverages off-the-shelf Vision-Language Models (VLMs) for autonomous navigation. Unlike most learning-based approaches that require extensive task- specific training and large-scale data collection, S2…

2025

3D Foundation Model-Based Loop Closing for Decentralized Collaborative SLAM

RA-L 2025

Decentralized Collaborative Simultaneous Localization and Mapping (C-SLAM) techniques often struggle to identify map overlaps due to significant viewpoint variations among robots. Motivated by recent advancements in 3D foundation models, which can register images despite large viewpoint differences,

Cited by 1SourceScholar
2025

Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models

NeurIPS 2025poster

Foundation models (FMs) trained with different objectives and data learn diverse representations, making some more effective than others for specific downstream tasks. Existing adaptation strategies, such as parameter-efficient fine-tuning, focus on individual models and do not exploit the complemen…

Cited by 0SourceScholar
2025

GraphSCENE: On-Demand Critical Scenario Generation for Autonomous Vehicles in Simulation

IROS 2025

Testing and validating Autonomous Vehicle (AV) performance in safety-critical and diverse scenarios is crucial before real-world deployment. However, manually creating such scenarios in simulation remains a significant and time-consuming challenge. This work introduces a novel method that generates

Cited by 2SourceScholar
2025

Select2Plan: Training-Free ICL-Based Planning Through VQA and Memory Retrieval

RA-L 2025

We introduce Select2Plan (S2P), a novel training-free framework for high-level robot planning that leverages off-the-shelf VLMs for autonomous navigation. Unlike most learning-based approaches that require extensive task-specific training and large-scale data collection, S2P overcomes the need for f

Cited by 4SourcecodeScholar
2025

Tiny LiDARs for Manipulator Self-Awareness: Sensor Characterization and Initial Localization Experiments

IROS 2025

For several tasks, ranging from manipulation to inspection, it is beneficial for robots to localize a target object in their surroundings. In this paper, we propose an approach that utilizes coarse point clouds obtained from miniaturized VL53L5CX Time-of-Flight (ToF) sensors (tiny LiDARs) to localiz

Cited by 3SourceScholar
2024

Masked γ-SSL: Learning Uncertainty Estimation via Masked Image Modeling

ICRA 2024poster

This work proposes a semantic segmentation network that produces high-quality uncertainty estimates in a single forward pass. We exploit general representations from foundation models and unlabelled datasets through a Masked Image Modeling (MIM) approach, which is robust to augmentation hyper-parame…

Cited by 1SourceScholar
2024

NeuralFloors++: Consistent Street-Level Scene Generation From BEV Semantic Maps

IROS 2024poster

Learning autonomous driving capabilities requires diverse and realistic training data. This has led to exploring generative techniques as an alternative to real-world data collection. In this paper we propose a method for synthesising photo-realistic urban driving scenes, along with semantic, instan…

Cited by 0SourceScholar
2024

NeuralFloors: Conditional Street-Level Scene Generation From BEV Semantic Maps via Neural Fields

RA-L 2024

Semantic Bird's Eye View (BEV) representations are a popular format, being easily interpretable and editable. However, synthesising ground-view images from BEVs is a difficult task as the system would need to learn both the mapping from BEV to Front View (FV) structure as well as to synthesise highl

Cited by 2SourceScholar
2024

That’s My Point: Compact Object-centric LiDAR Pose Estimation for Large-scale Outdoor Localisation

ICRA 2024poster

This paper is about 3D pose estimation on LiDAR scans with extremely minimal storage requirements to enable scalable mapping and localisation. We achieve this by clustering all points of segmented scans into semantic objects and representing them only with their respective centroid and semantic clas…

Cited by 3SourceScholar
2024

VDNA-PR: Using General Dataset Representations for Robust Sequential Visual Place Recognition

ICRA 2024poster

This paper adapts a general dataset representation technique to produce robust Visual Place Recognition (VPR) descriptors, crucial to enable real-world mobile robot localisation. Two parallel lines of work on VPR have shown, on one side, that general-purpose off-the-shelf feature representations can…

Cited by 1SourceScholar
2023

Explainable Action Prediction through Self-Supervision on Scene Graphs

ICRA 2023poster

This work explores scene graphs as a distilled representation of high-level information for autonomous driving, applied to future driver-action prediction. Given the scarcity and strong imbalance of data samples, we propose a self-supervision pipeline to infer representative and well-separated embed…

Cited by 13SourceScholar
2023

Self-Supervised Lidar Place Recognition in Overhead Imagery Using Unpaired Data

RSS 2023poster

As much as place recognition is crucial for navigation, mapping and collecting training ground truth, namely sensor data pairs across different locations, are costly and time-consuming. This paper tackles these by learning lidar place recognition on public overhead imagery and in a self-supervised…

2023

Visual DNA: Representing and Comparing Images Using Distributions of Neuron Activations

CVPR 2023poster

Selecting appropriate datasets is critical in modern computer vision. However, no general-purpose tools exist to evaluate the extent to which two datasets differ. For this, we propose representing images -- and by extension datasets -- using Distributions of Neuron Activations (DNAs). DNAs fit distr…

Cited by 13SourcePDFScholar
2023

Visual Servoing on Wheels: Robust Robot Orientation Estimation in Remote Viewpoint Control

IROS 2023poster

This work proposes a fast deployment pipeline for visually-servoed robots which does not assume anything about either the robot - e.g. sizes, colour or the presence of markers - or the deployment environment. Specifically, we apply a learning based approach to reliably estimate the pose of a robot i…

Cited by 5SourceScholar
2022

BoxGraph: Semantic Place Recognition and Pose Estimation from 3D LiDAR

IROS 2022poster

This paper is about extremely robust and lightweight localisation using LiDAR point clouds based on instance segmentation and graph matching. We model 3D point clouds as fully-connected graphs of semantically identified components where each vertex corresponds to an object instance and encodes its s…

Cited by 27SourceScholar
2022

Fast-MbyM: Leveraging Translational Invariance of the Fourier Transform for Efficient and Accurate Radar Odometry

ICRA 2022poster

Masking by Moving (MByM), provides robust and accurate radar odometry measurements through an exhaustive correlative search across discretised pose candidates. However, this dense search creates a significant computational bottleneck which hinders real-time performance when high-end GPUs are not ava…

Cited by 31SourceScholar
2022

What Goes Around: Leveraging a Constant-Curvature Motion Constraint in Radar Odometry

RA-L 2022

This letter presents a method that leverages vehicle motion constraints to refine data associations in a point-based radar odometry system. By using the strong prior on how a non-holonomic robot is constrained to move smoothly through its environment, we develop the necessary framework to estimate e

Cited by 18SourceScholar
2021

Fool Me Once: Robust Selective Segmentation via Out-of-Distribution Detection with Contrastive Learning

ICRA 2021poster

In this work, a neural network is trained to simultaneously perform segmentation and pixel-wise Out-of-Distribution (OoD) detection, such that the segmentation of unknown regions of scenes can be rejected. This is made possible by leveraging an OoD dataset with a novel contrastive objective and data…

Cited by 16SourceScholar
2021

Get to the Point: Learning Lidar Place Recognition and Metric Localisation Using Overhead Imagery

RSS 2021poster

This paper is about localising a robot in overhead images using lidar. Specifically; we show how to solve both place recognition and metric localisation of a lidar using only publicly available overhead imagery as a map proxy. This is in contrast to current approaches that rely on prior sensor map…

Cited by 29SourcePDFScholar
2020

Kidnapped Radar: Topological Radar Localisation using Rotationally-Invariant Metric Learning

ICRA 2020poster

This paper presents a system for robust, large-scale topological localisation using Frequency-Modulated Continuous-Wave scanning radar which extends the state-of-the-art by an efficient, learning-based approach to handle radar data for localisation. We learn a metric space for embedding polar radar…

Cited by 85SourceScholar
2020

Self-Supervised Localisation between Range Sensors and Overhead Imagery

RSS 2020poster

Publicly available satellite imagery can be an ubiquitous, cheap, and powerful tool for vehicle localisation when a prior sensor map is unavailable. However, satellite images are not directly comparable to data from ground range sensors because of their starkly different modalities. We present a l…

Cited by 26SourcePDFScholar
2019

Fast Radar Motion Estimation with a Learnt Focus of Attention using Weak Supervision

ICRA 2019poster

This paper is about fast motion estimation with scanning radar. We use weak supervision to train a focus of attention policy which actively down-samples the measurement stream before data association steps are undertaken. At training, we avoid laborious manual labelling by exploiting short-term sens…

Cited by 69SourceScholar