← Search

Alberto del Bimbo

12 accepted papers

2026

Mitigating Negative Flips via Margin Preserving Training

AAAI 2026technical

Minimizing inconsistencies across successive versions of an AI system is as crucial as reducing the overall error. In image classification, such inconsistencies manifest as negative flips, where an updated model misclassifies test samples that were previously classified correctly. This issue becomes

Cited by 0SourcePDFScholar
2025

$\boldsymbol{\lambda}$-Orthogonality Regularization for Compatible Representation Learning

NeurIPS 2025poster

Retrieval systems rely on representations learned by increasingly powerful models. However, due to the high training cost and inconsistencies in learned representations, there is significant interest in facilitating communication between representations and ensuring compatibility across independentl…

Cited by 0SourcecodeScholar
2024

Stationary Representations: Optimally Approximating Compatibility and Implications for Improved Model Replacements

CVPR 2024highlight

Learning compatible representations enables the interchangeable use of semantic features as models are updated over time. This is particularly relevant in search and retrieval systems where it is crucial to avoid reprocessing of the gallery images with the updated model. While recent research has sh…

2023

Zero-Shot Composed Image Retrieval with Textual Inversion

ICCV 2023poster

Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption that describes the difference between the two images. The high effort and cost required for labeling datasets for CIR hamper the widespread usage of existing methods,…

Cited by 126PDFcodeScholar
2022

Sparse to Dense Dynamic 3D Facial Expression Generation

CVPR 2022poster

In this paper, we propose a solution to the task of generating dynamic 3D facial expressions from a neutral 3D face and an expression label. This involves solving two sub-problems: (i) modeling the temporal dynamics of expressions, and (ii) deforming the neutral mesh to obtain the expressive counter…

Cited by 32PDFcodeScholar
2021

AdaVQA: Overcoming Language Priors with Adapted Margin Cosine Loss

IJCAI 2021poster

A number of studies point out that current Visual Question Answering (VQA) models are severely affected by the language prior problem, which refers to blindly making predictions based on the language shortcut. Some efforts have been devoted to overcoming this issue with delicate models. However, the…

2020

MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction

CVPR 2020poster

Autonomous vehicles are expected to drive in complex scenarios with several independent non cooperating agents. Path planning for safely navigating in such environments can not just rely on perceiving present location and motion of other agents. It requires instead to predict such variables in a far…

Cited by 153PDFcodeScholar
2020

Task-conditioned Domain Adaptation for Pedestrian Detection in Thermal Imagery

ECCV 2020poster

Pedestrian detection is a core problem in computer vision that sees broad application in video surveillance and, more recently, in advanced driving assistance systems. Despite its broad application and interest, it remains a challenging problem in part due to the vast range of conditions under which…

2018

Memory Based Online Learning of Deep Representations From Video Streams

CVPR 2018poster

We present a novel online unsupervised method for face identity learning from video streams. The method exploits deep face descriptors together with a memory based learning mechanism that takes advantage of the temporal coherence of visual data. Specifically, we introduce a discriminative descriptor…

Cited by 37SourcePDFScholar
2017

Deep Generative Adversarial Compression Artifact Removal

ICCV 2017poster

Compression artifacts arise in images whenever a lossy compression algorithm is applied. These artifacts eliminate details present in the original image, or add noise and small structures; because of these effects they make images less pleasant for the human eye, and may also lead to decreased perfo…

Cited by 257PDFScholar
2017

Group Re-Identification via Unsupervised Transfer of Sparse Features Encoding

ICCV 2017poster

Person re-identification is best known as the problem of associating a single person that is observed from one or more disjoint cameras. The existing literature has mainly addressed such an issue, neglecting the fact that people usually move in groups, like in crowded scenarios. We believe that the…

Cited by 66PDFScholar
2015

Representing 3D Texture on Mesh Manifolds for Retrieval and Recognition Applications

CVPR 2015poster

In this paper, we present and experiment a novel approach for representing texture of 3D mesh manifolds using local binary patterns (LBP). Using a recently proposed framework [37], we compute LBP directly on the mesh surface, either using geometric or photometric appearance. Compared to its depth-im…

Cited by 40SourcePDFScholar