← Search

Pierluigi Zama Ramirez

13 accepted papers

2026

Weight Space Representation Learning on Diverse NeRF Architectures

ICLR 2026poster

Neural Radiance Fields (NeRFs) have emerged as a groundbreaking paradigm for representing 3D objects and scenes by encoding shape and appearance information into the weights of a neural network. Recent studies have demonstrated that these weights can be used as input for frameworks designed to addre…

Cited by 0SourcecodeScholar
2025

SiM3D: Single-instance Multiview Multimodal and Multisetup 3D Anomaly Detection Benchmark

ICCV 2025poster

We propose SiM3D, the first benchmark considering the integration of multiview and multimodal information for comprehensive 3D anomaly detection and segmentation (ADS) where the task is to produce a voxel-based Anomaly Volume. Moreover, SiM3D focuses on a scenario of high interest in manufacturing:…

Cited by 0SourcePDFScholar
2025

Spatially-aware Weights Tokenization for NeRF-Language Models

NeurIPS 2025poster

Neural Radiance Fields (NeRFs) are neural networks -- typically multilayer perceptrons (MLPs) -- that represent the geometry and appearance of objects, with applications in vision, graphics, and robotics. Recent works propose understanding NeRFs with natural language using Multimodal Large Language…

Cited by 0SourceScholar
2024

Diffusion Models for Monocular Depth Estimation: Overcoming Challenging Conditions

ECCV 2024poster

"We present a novel approach designed to address the complexities posed by challenging, out-of-distribution data in the single-image depth estimation task. Starting with images that facilitate depth prediction due to the absence of unfavorable factors, we systematically generate new, user-defined sc…

2024

LLaNA: Large Language and NeRF Assistant

NeurIPS 2024poster

Multimodal Large Language Models (MLLMs) have demonstrated an excellent understanding of images and 3D data. However, both modalities have shortcomings in holistically capturing the appearance and geometry of objects. Meanwhile, Neural Radiance Fields (NeRFs), which encode information within the wei…

Cited by 4SourcePDFScholar
2024

Multimodal Industrial Anomaly Detection by Crossmodal Feature Mapping

CVPR 2024poster

Recent advancements have shown the potential of leveraging both point clouds and images to localize anomalies. Nevertheless their applicability in industrial manufacturing is often constrained by significant drawbacks such as the use of memory banks which leads to a substantial increase in terms of…

Cited by 23SourcePDFScholar
2024

Neural Processing of Tri-Plane Hybrid Neural Fields

ICLR 2024poster

Driven by the appealing properties of neural fields for storing and communicating 3D data, the problem of directly processing them to address tasks such as classification and part segmentation has emerged and has been investigated in recent works. Early approaches employ neural fields parameterized…

2023

Deep Learning on Implicit Neural Representations of Shapes

ICLR 2023poster

Implicit Neural Representations (INRs) have emerged in the last few years as a powerful tool to encode continuously a variety of different signals like images, videos, audio and 3D shapes. When applied to 3D shapes, INRs allow to overcome the fragmentation and shortcomings of the popular discrete r…

Cited by 56SourcePDFScholar
2023

Learning Depth Estimation for Transparent and Mirror Surfaces

ICCV 2023poster

Inferring the depth of transparent or mirror (ToM) surfaces represents a hard challenge for either sensors, algorithms, or deep networks. We propose a simple pipeline for learning to estimate depth properly for such surfaces with neural networks, without requiring any ground-truth annotation. We unv…

Cited by 43PDFScholar
2022

Open Challenges in Deep Stereo: The Booster Dataset

CVPR 2022poster

We present a novel high-resolution and challenging stereo dataset framing indoor scenes annotated with dense and accurate ground-truth disparities. Peculiar to our dataset is the presence of several specular and transparent surfaces, i.e. the main causes of failures for state-of-the-art stereo netwo…

Cited by 29PDFScholar
2022

RGB-Multispectral Matching: Dataset, Learning Methodology, Evaluation

CVPR 2022poster

We address the problem of registering synchronized color (RGB) and multi-spectral (MS) images featuring very different resolution by solving stereo matching correspondences. Purposely, we introduce a novel RGB-MS dataset framing 13 different scenes in indoor environments and providing a total of 34…

Cited by 11PDFScholar
2020

Distilled Semantics for Comprehensive Scene Understanding from Videos

CVPR 2020poster

Whole understanding of the surroundings is paramount to autonomous systems. Recent works have shown that deep neural networks can learn geometry (depth) and motion (optical flow) from a monocular video without any explicit supervision from ground truth annotations, particularly hard to source for th…

Cited by 91PDFcodeScholar