← Search

Fabio Cermelli

10 accepted papers

2025

Rethinking Cross-Modal Interaction for Efficient Referring Image Segmentation

RA-L 2025

Referring Image Segmentation, the task of finding and segmenting objects in an image conditioned on a natural language description, is crucial for human-robot collaboration. However, current RIS methods often implement visual-text alignment relying on computationally intensive Transformer-based self

Cited by 0SourceScholar
2024

PEM: Prototype-based Efficient MaskFormer for Image Segmentation

CVPR 2024poster

Recent transformer-based architectures have shown impressive results in the field of image segmentation. Thanks to their flexibility they obtain outstanding performance in multiple segmentation tasks such as semantic and panoptic under a single unified framework. To achieve such impressive performan…

2023

CoMFormer: Continual Learning in Semantic and Panoptic Segmentation

CVPR 2023poster

Continual learning for segmentation has recently seen increasing interest. However, all previous works focus on narrow semantic segmentation and disregard panoptic segmentation, an important task with real-world impacts. In this paper, we present the first continual learning model capable of operati…

2023

Unmasking Anomalies in Road-Scene Segmentation

ICCV 2023oral

Anomaly segmentation is a critical task for driving applications, and it is approached traditionally as a per-pixel classification problem. However, reasoning individually about each pixel without considering their contextual semantics results in high uncertainty around the objects' boundaries and n…

Cited by 45PDFcodeScholar
2022

FedDrive: Generalizing Federated Learning to Semantic Segmentation in Autonomous Driving

IROS 2022poster

Semantic Segmentation is essential to make self-driving vehicles autonomous, enabling them to understand their surroundings by assigning individual pixels to known categories. However, it operates on sensible data collected from the users' cars; thus, protecting the clients' privacy becomes a primar…

Cited by 68SourcecodeScholar
2022

Incremental Learning in Semantic Segmentation From Image Labels

CVPR 2022poster

Although existing semantic segmentation approaches achieve impressive results, they still struggle to update their models incrementally as new categories are uncovered. Furthermore, pixel-by-pixel annotations are expensive and time-consuming. This paper proposes a novel framework for Weakly Incremen…

Cited by 73PDFcodeScholar
2021

On the Challenges of Open World Recognition Under Shifting Visual Domains

RA-L 2021

Robotic visual systems operating in the wild must act in unconstrained scenarios, under different environmental conditions while facing a variety of semantic concepts, including unknown ones. To this end, recent works tried to empower visual object recognition methods with the capability to i) detec

Cited by 1SourcecodeScholar
2020

Boosting Deep Open World Recognition by Clustering

RA-L 2020

While convolutional neural networks have brought significant advances in robot vision, their ability is often limited to closed world scenarios, where the number of semantic concepts to be recognized is determined by the available training set. Since it is practically impossible to capture all possi

Cited by 25SourceScholar
2020

Modeling the Background for Incremental Learning in Semantic Segmentation

CVPR 2020poster

Despite their effectiveness in a wide range of tasks, deep architectures suffer from some important limitations. In particular, they are vulnerable to catastrophic forgetting, i.e. they perform poorly when they are required to update their model as new classes are available but the original training…

Cited by 378PDFcodeScholar
2019

The RGB-D Triathlon: Towards Agile Visual Toolboxes for Robots

IROS 2019poster

Deep networks have brought significant advances in robot perception, enabling to improve the capabilities of robots in several visual tasks, ranging from object detection and recognition to pose estimation, semantic scene segmentation and many others. Still, most approaches typically address visual…

Cited by 4SourcecodeScholar