← Search

Matteo Poggi

52 accepted papers

2026

Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo

CVPR 2026

Conventional frame-based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Event cameras offer an alternative visual representation with higher dynamic range free from such limitations. The complementary characteristics of the

Cited by 0SourcecodeScholar
2026

EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active Sensors

CVPR 2026

We propose EventHub, a novel framework for training deep-event stereo networks without ground truth annotations from costly active sensors, relying instead on standard color images. From these images, we derive either proxy annotations and proxy events through state-of-the-art novel view synthesis t

Cited by 0SourcecodeScholar
2026

FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM

AAAI 2026technical

We present FoundationSLAM, a learning-based monocular dense SLAM system that addresses the absence of geometric consistency in previous flow-based approaches for accurate and robust tracking and mapping. Our core idea is to bridge flow estimation with geometric reasoning by leveraging the guidance f

Cited by 0SourcePDFScholar
2026

Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic Segmentation

CVPR 2026

The effective integration and utilization of multimodal data acquired from image cameras and LiDAR is of paramount importance for perception systems. This paper proposes **I**mage-to-**P**oint Cloud **F**eature Back-**P**rojection (**IPFP**), a novel method for training multimodal fusion networks th

Cited by 0SourceScholar
2026

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

CVPR 2026

We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CLIP3R, a CLIP-informed 3D reconstruction module that predicts dense point maps from overlapping clips alongside object-le

Cited by 0SourceScholar
2025

Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation

ICCV 2025poster

Event cameras capture sparse, high-temporal-resolution visual information, making them particularly suitable for challenging environments with high-speed motion and strongly varying lighting conditions. However, the lack of large datasets with dense ground-truth depth annotations hinders learning-ba…

Cited by 0SourcePDFScholar
2025

Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian Splatting

NeurIPS 2025poster

We present Eve3D, a novel framework for dense 3D reconstruction based on 3D Gaussian Splatting (3DGS). While most existing methods rely on imperfect priors derived from pre-trained vision models, Eve3D fully leverages these priors by jointly optimizing both them and the 3DGS backbone. This joint opt…

Cited by 0SourceScholar
2025

HS-SLAM: Hybrid Representation with Structural Supervision for Improved Dense SLAM

ICRA 2025

NeRF-based SLAM has recently achieved promising results in tracking and reconstruction. However, existing methods face challenges in providing sufficient scene representation, capturing structural information, and maintaining global consistency in scenes emerging significant movement or being forgot

Cited by 5SourceScholar
2025

Learnable Fractional Reaction-Diffusion Dynamics for Under-Display ToF Imaging and Beyond

ICCV 2025poster

Under-display ToF imaging aims to achieve accurate depth sensing through a ToF camera placed beneath a screen panel. However, transparent OLED (TOLED) layers introduce severe degradations--such as signal attenuation, multi-path interference (MPI), and temporal noise--that significantly compromise de…

2025

Learning Temporally Consistent Video Depth from Video Diffusion Priors

CVPR 2025poster

This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate…

2025

Lightstereo: Channel Boost is All You Need for Efficient 2D Cost Aggregation

ICRA 2025

We present LightStereo, a cutting-edge stereomatching network crafted to accelerate the matching process. Departing from conventional methodologies that rely on aggregating computationally intensive 4D costs, LightStereo adopts the 3D cost volume as a lightweight alternative. While similar approache

Cited by 37SourcecodeScholar
2025

Self-supervised Monocular Depth Estimation for Dynamic Objects with Ground Propagation

IROS 2025

Self-supervised single-view depth estimation, trained on video sequences, faces significant challenges when dynamic objects are present in the training data, as they violate the basic multi-view geometry assumptions used to compute photometric losses. We propose a novel approach that leverages the r

Cited by 0SourcecodeScholar
2025

Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation

CVPR 2025poster

Open-vocabulary semantic segmentation models associate vision and text to label pixels from an undefined set of classes using textual queries, providing versatile performance on novel datasets. However, large shifts between training and test domains degrade their performance, requiring fine-tuning f…

2025

Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail

CVPR 2025poster

We introduce Stereo Anywhere, a novel stereo-matching framework that combines geometric constraints with robust priors from monocular depth Vision Foundation Models (VFMs). By elegantly coupling these complementary worlds through a dual-branch architecture, we seamlessly integrate stereo matching wi…

2025

ToF-Splatting: Dense SLAM using Sparse Time-of-Flight Depth and Multi-Frame Integration

ICCV 2025poster

Time-of-Flight (ToF) sensors provide efficient active depth sensing at relatively low power budgets; among such designs, only very sparse measurements from low-resolution sensors are considered to meet the increasingly limited power constraints of mobile and AR/VR devices. However, such extreme spar…

Cited by 0SourcePDFScholar
2024

Depth on Demand: Streaming Dense Depth from a Low Frame Rate Active Sensor

ECCV 2024poster

"High frame rate and accurate depth estimation plays an important role in several tasks crucial to robotics and automotive perception. To date, this can be achieved through ToF and LiDAR devices for indoor and outdoor applications, respectively. However, their applicability is limited by low frame r…

Cited by 3SourcePDFScholar
2024

Exploring Few-Beam LiDAR Assistance in Self-Supervised Multi-Frame Depth Estimation

IROS 2024poster

Self-supervised multi-frame depth estimation methods only require unlabeled monocular videos for training. However, most existing methods face challenges, including accuracy degradation caused by moving objects in dynamic scenes and scale ambiguity due to the absence of real-world references. In thi…

Cited by 0SourceScholar
2024

LiDAR-Event Stereo Fusion with Hallucinations

ECCV 2024poster

"Event stereo matching is an emerging technique to estimate depth from neuromorphic cameras; however, events are unlikely to trigger in the absence of motion or the presence of large, untextured regions, making the correspondence problem extremely challenging. Purposely, we propose integrating a ste…

2024

MaskingDepth: Masked Consistency Regularization for Semi-Supervised Monocular Depth Estimation

IROS 2024poster

We propose MaskingDepth, a semi-supervised learning framework for monocular depth estimation. MaskingDepth is designed to enforce consistency between the depths obtained from strongly-augmented images and the pseudo-depths derived from weakly-augmented images, which enables mitigating the reliance o…

Cited by 0SourcecodeScholar
2023

Active Stereo Without Pattern Projector

ICCV 2023poster

This paper proposes a novel framework integrating the principles of active stereo in standard passive camera systems without a physical pattern projector. We virtually project a pattern over the left and right images according to the sparse measurements obtained from a depth sensor. Any such devices…

Cited by 11PDFcodeScholar
2023

CompletionFormer: Depth Completion With Convolutions and Vision Transformers

CVPR 2023poster

Given sparse depths and the corresponding RGB images, depth completion aims at spatially propagating the sparse measurements throughout the whole image to get a dense depth prediction. Despite the tremendous progress of deep-learning-based depth completion methods, the locality of the convolutional…

2023

GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction

ICCV 2023poster

Neural implicit representations have recently demonstrated compelling results on dense Simultaneous Localization And Mapping (SLAM) but suffer from the accumulation of errors in camera tracking and distortion in the reconstruction. Purposely, we present GO-SLAM, a deep-learning-based dense visual SL…

Cited by 135PDFcodeScholar
2023

GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor Scenes

ICCV 2023poster

This paper tackles the challenges of self-supervised monocular depth estimation in indoor scenes caused by large rotation between frames and low texture. We ease the learning process by obtaining coarse camera poses from monocular sequences through multi-view geometry to deal with the former. Howeve…

Cited by 25PDFcodeScholar
2023

Learning Depth Estimation for Transparent and Mirror Surfaces

ICCV 2023poster

Inferring the depth of transparent or mirror (ToM) surfaces represents a hard challenge for either sensors, algorithms, or deep networks. We propose a simple pipeline for learning to estimate depth properly for such surfaces with neural networks, without requiring any ground-truth annotation. We unv…

Cited by 43PDFScholar
2023

TemporalStereo: Efficient Spatial-Temporal Stereo Matching Network

IROS 2023poster

We present TemporalStereo, a coarse-to-fine stereo matching network that is highly efficient, and able to effectively exploit the past geometry and context information to boost matching accuracy. Our network leverages sparse cost volume and proves to be effective when a single stereo pair is given.…

Cited by 14SourcecodeScholar
2023

To Adapt or Not to Adapt? Real-Time Adaptation for Semantic Segmentation

ICCV 2023poster

The goal of Online Domain Adaptation for semantic segmentation is to handle unforeseeable domain changes that occur during deployment, like sudden weather events. However, the high computational costs associated with brute-force adaptation make this paradigm unfeasible for real-world applications. I…

Cited by 13PDFcodeScholar
2022

Meta-confidence estimation for stereo matching

ICRA 2022poster

We propose a novel framework to estimate the confidence of a disparity map taking into account, for the first time, the uncertainty affecting the confidence estimation process itself. Conversely to other tasks such as disparity estimation, the uncertainty of confidence directly hints that the confid…

Cited by 2SourceScholar
2022

Online Domain Adaptation for Semantic Segmentation in Ever-Changing Conditions

ECCV 2022poster

"Unsupervised Domain Adaptation (UDA) aims at reducing the domain gap between training and testing data and is, in most cases, carried out in offline manner. However, domain changes may occur continuously and unpredictably during deployment (e.g. sudden weather changes). In such conditions, deep neu…

2022

Open Challenges in Deep Stereo: The Booster Dataset

CVPR 2022poster

We present a novel high-resolution and challenging stereo dataset framing indoor scenes annotated with dense and accurate ground-truth disparities. Peculiar to our dataset is the presence of several specular and transparent surfaces, i.e. the main causes of failures for state-of-the-art stereo netwo…

Cited by 29PDFScholar
2022

RGB-Multispectral Matching: Dataset, Learning Methodology, Evaluation

CVPR 2022poster

We address the problem of registering synchronized color (RGB) and multi-spectral (MS) images featuring very different resolution by solving stereo matching correspondences. Purposely, we introduce a novel RGB-MS dataset framing 13 different scenes in indoor environments and providing a total of 34…

Cited by 11PDFScholar
2022

Unsupervised confidence for LiDAR depth maps and applications

IROS 2022poster

Depth perception is pivotal in many fields, such as robotics and autonomous driving, to name a few. Consequently, depth sensors such as LiDARs rapidly spread in many applications. The 3D point clouds generated by these sensors must often be coupled with an RGB camera to understand the framed scene s…

Cited by 13SourcecodeScholar
2020

Distilled Semantics for Comprehensive Scene Understanding from Videos

CVPR 2020poster

Whole understanding of the surroundings is paramount to autonomous systems. Recent works have shown that deep neural networks can learn geometry (depth) and motion (optical flow) from a monocular video without any explicit supervision from ground truth annotations, particularly hard to source for th…

Cited by 91PDFcodeScholar
2020

On the Uncertainty of Self-Supervised Monocular Depth Estimation

CVPR 2020poster

Self-supervised paradigms for monocular depth estimation are very appealing since they do not require ground truth annotations at all. Despite the astonishing results yielded by such methodologies, learning to reason about the uncertainty of the estimated depth maps is of paramount importance for pr…

Cited by 309PDFcodeScholar
2020

Real-Time Semantic Stereo Matching

ICRA 2020poster

Scene understanding is paramount in robotics, self-navigation, augmented reality, and many other fields. To fully accomplish this task, an autonomous agent has to infer the 3D structure of the sensed scene (to know where it looks at) and its content (to know what it sees). To tackle the two tasks, d…

Cited by 84SourceScholar
2020

Reversing the cycle: self-supervised deep stereo through enhanced monocular distillation

ECCV 2020poster

In many fields, self-supervised learning solutions are rapidly evolving and filling the gap with supervised approaches. This fact occurs for depth estimation based on either monocular or stereo, with the latter often providing a valid source of self-supervision for the former. In contrast, to soften…

2020

Self-adapting confidence estimation for stereo

ECCV 2020poster

Estimating the confidence of disparity maps inferred by a stereo algorithm has become a very relevant task in the years, due to the increasing number of applications leveraging such cue. Although self-supervised learning has recently spread across many computer vision tasks, it has been barely consi…

2019

Learning Monocular Depth Estimation Infusing Traditional Stereo Knowledge

CVPR 2019poster

Depth estimation from a single image represents a fascinating, yet challenging problem with countless applications. Recent works proved that this task could be learned without direct supervision from ground truth labels leveraging image synthesis on sequences or stereo pairs. Focusing on this second…

Cited by 278PDFcodeScholar
2018

Beyond local reasoning for stereo confidence estimation with deep learning

ECCV 2018poster

Confidence measures for stereo gained popularity in recent years due to their improved capability to detect outliers and the increasing number of applications exploiting these cues. In this field, convolutional neural networks achieved top-performance compared to other known techniques in the litera…

2018

Towards Real-Time Unsupervised Monocular Depth Estimation on CPU

IROS 2018poster

Unsupervised depth estimation from a single image is a very attractive technique with several implications in robotic, autonomous navigation, augmented reality and so on. This topic represents a very challenging task and the advent of deep learning enabled to tackle this problem with excellent resul…

Cited by 189SourcecodeScholar