← Search

Andrea Cavallaro

34 accepted papers

2026

Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training

ICASSP 2026poster

Diffusion-based generative models have achieved state-of-the-art performance for perceptual quality in speech enhancement (SE). However, their iterative nature requires numerous Neural Function Evaluations (NFEs), posing a challenge for real-time applications. On the contrary, flow matching offers a…

Cited by 0SourcePDFScholar
2025

NaviFormer: A Deep Reinforcement Learning Transformer-like Model to Holistically Solve the Navigation Problem

IROS 2025

Path planning is usually solved by addressing either the (high-level) route planning problem (waypoint sequencing to achieve the final goal) or the (low-level) path planning problem (trajectory prediction between two waypoints avoiding collisions). However, real-world problems usually require simult

Cited by 0SourceScholar
2024

Open-Vocabulary Object 6D Pose Estimation

CVPR 2024highlight

We introduce the new setting of open-vocabulary object 6D pose estimation in which a textual prompt is used to specify the object of interest. In contrast to existing approaches in our setting (i) the object of interest is specified solely through the textual prompt (ii) no object model (e.g. CAD or…

2022

Audio-Visual Object Classification for Human-Robot Collaboration

ICASSP 2022accepted

Human-robot collaboration requires the contactless estimation of the physical properties of containers manipulated by a person, for example while pouring content in a cup or moving a food box. Acoustic and visual signals can be used to estimate the physical properties of such objects, which may vary…

Cited by 0SourceScholar
2022

Is Cross-Attention Preferable to Self-Attention for Multi-Modal Emotion Recognition?

ICASSP 2022accepted

Humans express their emotions via facial expressions, voice intonation and word choices. To infer the nature of the underlying emotion, recognition models may use a single modality, such as vision, audio, and text, or a combination of modalities. Generally, models that fuse complementary information…

Cited by 0SourceScholar
2022

Training Privacy-Preserving Video Analytics Pipelines by Suppressing Features That Reveal Information About Private Attributes

ICASSP 2022accepted

Deep neural networks are increasingly deployed for scene analytics, including to evaluate the attention and reaction of people exposed to out-of-home advertisements. However, the features extracted by a deep neural network that was trained to predict a specific, consensual attribute (e.g. emotion) m…

Cited by 0SourceScholar
2021

FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial Disturbances

ICASSP 2021accepted

Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates imperceptible adversarial perturbations against a speaker i…

Cited by 0SourceScholar
2021

Robust Latent Representations Via Cross-Modal Translation and Alignment

ICASSP 2021accepted

Multi-modal learning relates information across observation modalities of the same physical phenomenon to leverage complementary information. Most multi-modal machine learning methods require that all the modalities used for training are also available for testing. This is a limitation when signals…

Cited by 0SourceScholar
2020

Benchmark for Human-to-Robot Handovers of Unseen Containers With Unknown Filling

RA-L 2020

The real-time estimation through vision of the physical properties of objects manipulated by humans is important to inform the control of robots for performing accurate and safe grasps of objects handed over by humans. However, estimating the 3D pose and dimensions of previously unseen objects using

Cited by 43SourceScholar
2020

Multi-View Shape Estimation of Transparent Containers

ICASSP 2020accepted

The 3D localisation of an object and the estimation of its properties, such as shape and dimensions, are challenging under varying degrees of transparency and lighting conditions. In this paper, we propose a method for jointly localising container-like objects and estimating their dimensions using t…

Cited by 0SourceScholar
2019

Accurate Target Annotation in 3D from Multimodal Streams

ICASSP 2019accepted

Accurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverag…

Cited by 0SourceScholar
2019

Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement

IROS 2019poster

We present an audio-visual dataset recorded outdoors from a quadcopter and discuss baseline results for multiple applications. The dataset includes a scenario for source localization and sound enhancement with up to two static sources, and a scenario for source localization and tracking with a movin…

Cited by 30SourceScholar
2019

Omni-Scale Feature Learning for Person Re-Identification

ICCV 2019poster

As an instance-level recognition problem, person re-identification (ReID) relies on discriminative features, which not only capture different spatial scales but also encapsulate an arbitrary combination of multiple scales. We callse features of both homogeneous and heterogeneous scales omni-scale fe…

Cited by 1039PDFcodeScholar
2019

Scene Privacy Protection

ICASSP 2019accepted

Images shared on social media are routinely analysed by classifiers for content annotation and user profiling. These automatic inferences reveal to the service provider sensitive information that a naive user might want to keep private. To address this problem, we present a method designed to distor…

Cited by 0SourceScholar
2018

3D Mouth Tracking from a Compact Microphone Array Co-Located with a camera

ICASSP 2018accepted

We address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood c…

Cited by 0SourceScholar
2018

a Multi-Perspective Approach to Anomaly Detection for Self -Aware Embodied Agents

ICASSP 2018accepted

This paper focuses on multi-sensor anomaly detection for moving cognitive agents using both external and private first-person visual observations. Both observation types are used to characterize agents motion in a given environment. The proposed method generates locally uniform motion models by divi…

Cited by 0SourceScholar
2017

3D audio-visual speaker tracking with an adaptive particle filter

ICASSP 2017accepted

We propose an audio-visual fusion algorithm for 3D speaker tracking from a localised multi-modal sensor platform composed of a camera and a small microphone array. After extracting audio-visual cues from individual modalities we fuse them adaptively using their reliability in a particle filter frame…

Cited by 0SourceScholar
2016

Rate-adaptive multicast video streaming from teams of micro aerial vehicles

ICRA 2016

Video multicasting from cameras mounted on micro aerial vehicles (MAVs) is desirable for applications such as search and rescue, surveillance and disaster management. Because of the mobility of the video sources and the high data-rate of videos, the transmission rate should be adapted to the task at

Cited by 8SourceScholar
2015

Multiscale observation of multiple moving targets using Micro Aerial Vehicles

IROS 2015poster

This paper presents a centralized algorithm for multi-scale observation of multiple moving targets using a team of Micro Aerial Vehicles (MAVs). The proposed algorithm is appropriate when MAVs can observe targets at different elevations with the objective of jointly maximizing duration and resolutio…

Cited by 31SourceScholar