← Search

Luigi Di Stefano

24 accepted papers

2026

Weight Space Representation Learning on Diverse NeRF Architectures

ICLR 2026poster

Neural Radiance Fields (NeRFs) have emerged as a groundbreaking paradigm for representing 3D objects and scenes by encoding shape and appearance information into the weights of a neural network. Recent studies have demonstrated that these weights can be used as input for frameworks designed to addre…

Cited by 0SourcecodeScholar
2025

SiM3D: Single-instance Multiview Multimodal and Multisetup 3D Anomaly Detection Benchmark

ICCV 2025poster

We propose SiM3D, the first benchmark considering the integration of multiview and multimodal information for comprehensive 3D anomaly detection and segmentation (ADS) where the task is to produce a voxel-based Anomaly Volume. Moreover, SiM3D focuses on a scenario of high interest in manufacturing:…

Cited by 0SourcePDFScholar
2025

Spatially-aware Weights Tokenization for NeRF-Language Models

NeurIPS 2025poster

Neural Radiance Fields (NeRFs) are neural networks -- typically multilayer perceptrons (MLPs) -- that represent the geometry and appearance of objects, with applications in vision, graphics, and robotics. Recent works propose understanding NeRFs with natural language using Multimodal Large Language…

Cited by 0SourceScholar
2024

LLaNA: Large Language and NeRF Assistant

NeurIPS 2024poster

Multimodal Large Language Models (MLLMs) have demonstrated an excellent understanding of images and 3D data. However, both modalities have shortcomings in holistically capturing the appearance and geometry of objects. Meanwhile, Neural Radiance Fields (NeRFs), which encode information within the wei…

Cited by 4SourcePDFScholar
2024

Multimodal Industrial Anomaly Detection by Crossmodal Feature Mapping

CVPR 2024poster

Recent advancements have shown the potential of leveraging both point clouds and images to localize anomalies. Nevertheless their applicability in industrial manufacturing is often constrained by significant drawbacks such as the use of memory banks which leads to a substantial increase in terms of…

Cited by 23SourcePDFScholar
2024

Neural Processing of Tri-Plane Hybrid Neural Fields

ICLR 2024poster

Driven by the appealing properties of neural fields for storing and communicating 3D data, the problem of directly processing them to address tasks such as classification and part segmentation has emerged and has been investigated in recent works. Early approaches employ neural fields parameterized…

2023

Deep Learning on Implicit Neural Representations of Shapes

ICLR 2023poster

Implicit Neural Representations (INRs) have emerged in the last few years as a powerful tool to encode continuously a variety of different signals like images, videos, audio and 3D shapes. When applied to 3D shapes, INRs allow to overcome the fragmentation and shortcomings of the popular discrete r…

Cited by 56SourcePDFScholar
2023

Learning Depth Estimation for Transparent and Mirror Surfaces

ICCV 2023poster

Inferring the depth of transparent or mirror (ToM) surfaces represents a hard challenge for either sensors, algorithms, or deep networks. We propose a simple pipeline for learning to estimate depth properly for such surfaces with neural networks, without requiring any ground-truth annotation. We unv…

Cited by 43PDFScholar
2023

ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World Objects

CVPR 2023highlight

In this paper, we focus on the problem of rendering novel views from a Neural Radiance Field (NeRF) under unobserved light conditions. To this end, we introduce a novel dataset, dubbed ReNe (Relighting NeRF), framing real world objects under one-light-at-time (OLAT) conditions, annotated with accura…

2022

Open Challenges in Deep Stereo: The Booster Dataset

CVPR 2022poster

We present a novel high-resolution and challenging stereo dataset framing indoor scenes annotated with dense and accurate ground-truth disparities. Peculiar to our dataset is the presence of several specular and transparent surfaces, i.e. the main causes of failures for state-of-the-art stereo netwo…

Cited by 29PDFScholar
2022

RGB-Multispectral Matching: Dataset, Learning Methodology, Evaluation

CVPR 2022poster

We address the problem of registering synchronized color (RGB) and multi-spectral (MS) images featuring very different resolution by solving stereo matching correspondences. Purposely, we introduce a novel RGB-MS dataset framing 13 different scenes in indoor environments and providing a total of 34…

Cited by 11PDFScholar
2020

Ambiguity in Sequential Data: Predicting Uncertain Futures With Recurrent Models

RA-L 2020

Ambiguity is inherently present in many machine learning tasks, but especially for sequential models seldom accounted for, as most only output a single prediction. In this work we propose an extension of the Multiple Hypothesis Prediction (MHP) model to handle ambiguous predictions with sequential d

Cited by 7SourceScholar
2020

Distilled Semantics for Comprehensive Scene Understanding from Videos

CVPR 2020poster

Whole understanding of the surroundings is paramount to autonomous systems. Recent works have shown that deep neural networks can learn geometry (depth) and motion (optical flow) from a monocular video without any explicit supervision from ground truth annotations, particularly hard to source for th…

Cited by 91PDFcodeScholar
2020

Learning to Orient Surfaces by Self-supervised Spherical CNNs

NeurIPS 2020poster

Defining and reliably finding a canonical orientation for 3D surfaces is key to many Computer Vision and Robotics applications. This task is commonly addressed by handcrafted algorithms exploiting geometric cues deemed as distinctive and robust by the designer. Yet, one might conjecture that humans…

2019

GFrames: Gradient-Based Local Reference Frame for 3D Shape Matching

CVPR 2019oral

We introduce GFrames, a novel local reference frame (LRF) construction for 3D meshes and point clouds. GFrames are based on the computation of the intrinsic gradient of a scalar field defined on top of the input shape. The resulting tangent vector field defines a repeatable tangent direction of the…

Cited by 34PDFScholar
2019

Learning to Adapt for Stereo

CVPR 2019poster

Real world applications of stereo depth estimation require models that are robust to dynamic variations in the environment. Even though deep learning based stereo methods are successful, they often fail to generalize to unseen variations in the environment, making them less suitable for practical ap…

Cited by 93PDFcodeScholar
2017

On-The-Fly Adaptation of Regression Forests for Online Camera Relocalisation

CVPR 2017oral

Camera relocalisation is an important problem in computer vision, with applications in simultaneous localisation and mapping, virtual/augmented reality and navigation. Common techniques either match the current image against keyframes with known poses coming from a tracker, or establish 2D-to-3D cor…

Cited by 141PDFScholar
2015

Large-Scale and Drift-Free Surface Reconstruction Using Online Subvolume Registration

CVPR 2015poster

Depth cameras have helped commoditize 3D digitization of the real-world. It is now feasible to use a single Kinect-like camera to scan in an entire building or other large-scale scenes. At large scales, however, there is an inherent challenge of dealing with distortions and drift due to accumulated…

Cited by 80SourcePDFScholar
2015

Learning a Descriptor-Specific 3D Keypoint Detector

ICCV 2015poster

Keypoint detection represents the first stage in the majority of modern computer vision pipelines based on automatically established correspondences between local descriptors. However, no standard solution has emerged yet in the case of 3D data such as point clouds or meshes, which exhibit high vari…

Cited by 54PDFScholar