← Search

Francesco Ferroni

11 accepted papers

2026

DuoGen: Towards Autonomous Interleaved Multimodal Generation

CVPR 2026

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved generation models under general instructions remains limited

Cited by 0SourceScholar
2024

Better Call SAL: Towards Learning to Segment Anything in Lidar

ECCV 2024poster

"We propose the (Segment Anything in Lidar) method consisting of a text-promptable zero-shot model for segmenting and classifying any object in Lidar, and a pseudo-labeling engine that facilitates model training without manual supervision. While the established paradigm for (LPS) relies on manual su…

2024

LCA-on-the-Line: Benchmarking Out of Distribution Generalization with Class Taxonomies

ICML 2024oral

We tackle the challenge of predicting models' Out-of-Distribution (OOD) performance using in-distribution (ID) measurements without requiring OOD data. Existing evaluations with ``Effective robustness'', which use ID accuracy as an indicator of OOD accuracy, encounter limitations when models are tra…

2024

SeMoLi: What Moves Together Belongs Together

CVPR 2024poster

We tackle semi-supervised object detection based on motion cues. Recent results suggest that heuristic-based clustering methods in conjunction with object trackers can be used to pseudo-label instances of moving objects and use these as supervisory signals to train 3D object detectors in Lidar data…

Cited by 8SourcePDFScholar
2023

Lidar Panoptic Segmentation and Tracking without Bells and Whistles

IROS 2023poster

State-of-the-art lidar panoptic segmentation (LPS) methods follow “bottom-up” segmentation-centric fashion wherein they build upon semantic segmentation networks by utilizing clustering to obtain object instances. In this paper, we re-think this approach and propose a surprisingly simple yet effecti…

Cited by 8SourcecodeScholar
2023

Pix2map: Cross-Modal Retrieval for Inferring Street Maps From Images

CVPR 2023poster

Self-driving vehicles rely on urban street maps for autonomous navigation. In this paper, we introduce Pix2Map, a method for inferring urban street map topology directly from ego-view images, as needed to continually update and expand existing maps. This is a challenging task, as we need to infer a…

Cited by 10SourcePDFScholar
2020

Learning Flat Latent Manifolds with VAEs

ICML 2020poster

Measuring the similarity between data points often requires domain knowledge, which can in parts be compensated by relying on unsupervised methods such as latent-variable models, where similarity/distance is estimated in a more compact latent space. Prevalent is the use of the Euclidean metric, whic…

Cited by 53SourcePDFScholar
2019

Sampling-Free Epistemic Uncertainty Estimation Using Approximated Variance Propagation

ICCV 2019oral

We present a sampling-free approach for computing the epistemic uncertainty of a neural network. Epistemic uncertainty is an important quantity for the deployment of deep neural networks in safety-critical applications, since it represents how much one can trust predictions on new data. Recently pro…

Cited by 183PDFcodeScholar