← Search

Stefano Gasperini

12 accepted papers

2026

Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion

ICML 2026poster

Forecasting the evolution of dynamic environments is crucial for autonomous agents. While generative world models have recently achieved high photorealism in 2D video synthesis, by mixing within the image plane ego-motion and environmental dynamics, they exhibit physical inconsistencies, such as mor…

Cited by 0SourceScholar
2026

InstantPose: Zero-Shot Instance-Level 6D Pose Estimation from a Single View

ICRA 2026poster

Object pose estimation using visual data is crucial for robotic interaction with the environment. Many existing instance-level methods are restricted by their requirements for 3D CAD models or multiple object views, which limits their flexibility and generalizability. Overcoming this limitation is c…

Cited by 0SourceScholar
2025

InstantPose: Zero-Shot Instance-Level 6D Pose Estimation From a Single View

RA-L 2025

Object pose estimation using visual data is crucial for robotic interaction with the environment. Many existing instance-level methods are restricted by their requirements for 3D CAD models or multiple object views, which limits their flexibility and generalizability. Overcoming this limitation is c

Cited by 4SourceScholar
2025

Prior2Former - Evidential Modeling of Mask Transformers for Assumption-Free Open-World Panoptic Segmentation

ICCV 2025poster

In panoptic segmentation, individual instances must be separated within semantic classes. As state-of-the-art methods rely on a pre-defined set of classes, they struggle with novel categories and out-of-distribution (OOD) data. This is particularly problematic in safety-critical applications, such a…

Cited by 0SourcePDFScholar
2025

VoxNeRF: Bridging Voxel Representation and Neural Radiance Fields for Enhanced Indoor View Synthesis

RA-L 2025

The generation of high-fidelity view synthesis is essential for robotic navigation and interaction but remains challenging, particularly in indoor environments and real-time scenarios. Existing techniques often require significant computational resources for both training and rendering, and they fre

Cited by 2SourceScholar
2024

Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation

IROS 2024

Estimating the pose of objects through vision is essential to make robotic platforms interact with the environment. Yet, it presents many challenges, often related to the lack of flexibility and generalizability of state-of-the-art solutions. Diffusion models are a cutting-edge neural architecture t

Cited by 13SourceScholar
2023

Robust Monocular Depth Estimation under Challenging Conditions

ICCV 2023accepted

While state-of-the-art monocular depth estimation approaches achieve impressive results in ideal settings, they are highly unreliable under challenging illumination and weather conditions, such as at nighttime or in the presence of rain. In this paper, we uncover these safety-critical issues and tac…

2023

Segmenting Known Objects and Unseen Unknowns without Prior Knowledge

ICCV 2023poster

Panoptic segmentation methods assign a known class to each pixel given in input. Even for state-of-the-art approaches, this inevitably enforces decisions that systematically lead to wrong predictions for objects outside the training categories. However, robustness against out-of-distribution samples…

Cited by 10PDFcodeScholar
2022

3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection

CVPR 2022poster

As 3D object detection on point clouds relies on the geometrical relationships between the points, non-standard object shapes can hinder a method's detection capability. However, in safety-critical settings, robustness to out-of-domain and long-tail samples is fundamental to circumvent dangerous iss…

Cited by 69PDFcodeScholar
2022

CertainNet: Sampling-Free Uncertainty Estimation for Object Detection

RA-L 2022

Estimating the uncertainty of a neural network plays a fundamental role in safety-critical settings. In perception for autonomous driving, measuring the uncertainty means providing additional calibrated information to downstream tasks, such as path planning, that can use it towards safe navigation.

Cited by 29SourceScholar
2021

Panoster: End-to-End Panoptic Segmentation of LiDAR Point Clouds

RA-L 2021

Panoptic segmentation has recently unified semantic and instance segmentation, previously addressed separately, thus taking a step further towards creating more comprehensive and efficient perception systems. In this letter, we present Panoster, a novel proposal-free panoptic segmentation method for

Cited by 76SourceScholar
2020

Signal Clustering With Class-Independent Segmentation

ICASSP 2020accepted

Radar signals have been dramatically increasing in complexity, limiting the source separation ability of traditional approaches. In this paper we propose a Deep Learning-based clustering method, which encodes concurrent signals into images, and, for the first time, tackles clustering with image segm…

Cited by 0SourceScholar