← Search

Vittorio Murino

21 accepted papers

2026

Distribution Alignment for One-Shot Federated Learning via Optimal Transport

ICML 2026poster

One-Shot Federated Learning (OSFL) addresses extreme communication regimes in which clients interact with the server only once, amplifying the impact of heterogeneous client data distributions. In particular, the interaction of domain shift and label shift across clients induces misaligned feature r…

Cited by 0SourceScholar
2026

Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment

CVPR 2026

We introduce a lifelong imitation learning framework that enables continual policy refinement across sequential tasks under realistic memory and data constraints. Our approach departs from conventional experience replay by operating entirely in a multimodal latent space, where compact representation

Cited by 0SourcecodeScholar
2025

Diffusing DeBias: Synthetic Bias Amplification for Model Debiasing

NeurIPS 2025poster

The effectiveness of deep learning models in classification tasks is often challenged by the quality and quantity of training data whenever they are affected by strong spurious correlations between specific attributes and target labels. This results in a form of bias affecting training data, which t…

Cited by 0SourcecodeScholar
2023

Audio-Visual Inpainting: Reconstructing Missing Visual Information with Sound

ICASSP 2023accepted

We tackle audio-visual inpainting, the problem of completing an image in such a way to be consistent with the sound associated to the scene. To this end, we propose a multimodal, audio-visual inpainting method (AVIN), and show how to leverage sound to reconstruct semantically consistent images. AVIN…

Cited by 0SourceScholar
2021

Audio-Visual Localization by Synthetic Acoustic Image Generation

AAAI 2021technical

Acoustic images constitute an emergent data modality for multimodal scene understanding. Such images have the peculiarity to distinguish the spectral signature of sounds coming from different directions in space, thus providing richer information than the one derived from mono and binaural microphon…

2020

Leveraging Acoustic Images for Effective Self-Supervised Audio Representation Learning

ECCV 2020poster

In this paper, we propose the use of a new modality characterized by a richer information content, namely acoustic images, for the sake of audio-visual scene understanding. Each pixel in such images is characterized by a spectral signature, associated to a specific direction in space and obtained by…

2019

Addressing Model Vulnerability to Distributional Shifts Over Image Transformation Sets

ICCV 2019poster

We are concerned with the vulnerability of computer vision models to distributional shifts. We formulate a combinatorial optimization problem that allows evaluating the regions in the image space where a given model is more vulnerable, in terms of image transformations applied to the input, and face…

Cited by 139PDFcodeScholar
2018

Adversarial Feature Augmentation for Unsupervised Domain Adaptation

CVPR 2018poster

Recent works showed that Generative Adversarial Networks (GANs) can be successfully applied in unsupervised domain adaptation, where, given a labeled source dataset and an unlabeled target dataset, the goal is to train powerful classifiers for the target samples. In particular, it was shown that a G…

Cited by 303SourcePDFScholar
2018

Dropout as a Low-Rank Regularizer for Matrix Factorization

AISTATS 2018poster

Regularization for matrix factorization (MF) and approximation problems has been carried out in many different ways. Due to its popularity in deep learning, dropout has been applied also for this class of problems. Despite its solid empirical performance, the theoretical properties of dropout as a r…

Cited by 0SourcePDFScholar
2018

Excitation Backprop for RNNs

CVPR 2018poster

Deep models are state-of-the-art or many vision tasks including video action recognition and video captioning. Models are trained to caption or classify activity in videos, but little is known about the evidence used to make such decisions. Grounding decisions made by deep networks has been studied…

2018

Generalizing to Unseen Domains via Adversarial Data Augmentation

NeurIPS 2018poster

We are concerned with learning models that generalize well to different unseen domains. We consider a worst-case formulation over data distributions that are near the source domain in the feature space. Only using training data from a single source distribution, we propose an iterative procedure tha…

2018

Minimal-Entropy Correlation Alignment for Unsupervised Deep Domain Adaptation

ICLR 2018poster

In this work, we face the problem of unsupervised domain adaptation with a novel deep learning approach which leverages our finding that entropy minimization is induced by the optimal alignment of second order statistics between source and target domains. We formally demonstrate this hypothesis and,…

Cited by 204SourcePDFScholar
2018

Modality Distillation with Multiple Stream Networks for Action Recognition

ECCV 2018poster

Diverse input data modalities can provide complementary cues for several tasks, usually leading to more robust algorithms and better performance. However, while a (training) dataset could be accurately designed to include a variety of sensory inputs, it is often the case that not all modalities are…

2017

Summarization and Classification of Wearable Camera Streams by Learning the Distributions Over Deep Features of Out-Of-Sample Image Sequences

ICCV 2017poster

A popular approach to training classifiers of new image classes is to use lower levels of a pre-trained feed-forward neural network and retrain only the top. Thus, most layers simply serve as highly nonlinear feature extractors. While these features were found useful for classifying a variety of sce…

Cited by 8PDFScholar
2017

Unsupervised Adaptive Re-Identification in Open World Dynamic Camera Networks

CVPR 2017spotlight

Person re-identification is an open and challenging problem in computer vision. Existing approaches have concentrated on either designing the best feature representation or learning optimal matching metrics in a static setting where the number of cameras are fixed in a network. Most approaches have…

Cited by 48PDFScholar
2016

Approximate Log-Hilbert-Schmidt Distances Between Covariance Operators for Image Classification

CVPR 2016poster

This paper presents a novel framework for visual object recognition using infinite-dimensional covariance operators of input features, in the paradigm of kernel methods on infinite-dimensional Riemannian manifolds. Our formulation provides a rich representation of image features by exploiting their…

Cited by 24PDFScholar
2016

Fast 6D pose estimation for texture-less objects from a single RGB image

ICRA 2016

A fundamental step to solve bin-picking and grasping problems is the accurate estimation of an object 3D pose. Such visual task usually rely on profusely textured objects: standard procedures such as detection of interest points or computation of appearance-based descriptors are favoured by using a

Cited by 42SourceScholar
2015

Learning With Dataset Bias in Latent Subcategory Models

CVPR 2015poster

Latent subcategory models (LSMs) offer significant improvements over training flat classifiers such as linear SVMs. Training LSMs is a challenging task due to the potentially large number of local optima in the objective function and the increased model complexity which requires large training set s…

Cited by 17SourcePDFScholar
2015

Sparse Representation Classification With Manifold Constraints Transfer

CVPR 2015poster

The fact that image data samples lie on a manifold has been successfully exploited in many learning and inference problems. In this paper we leverage the specific structure of data in order to improve recognition accuracies in general recognition tasks. In particular we propose a novel framework tha…

Cited by 66SourcePDFScholar