← Search

Stefano Tubaro

23 accepted papers

2026

MULTI-TASK TRANSFORMER FOR EXPLAINABLE SPEECH DEEPFAKE DETECTION VIA FORMANT MODELING

ICASSP 2026poster

In this work, we introduce a multi-task transformer for speech deepfake detection, capable of predicting formant trajectories and voicing patterns over time, ultimately classifying speech as real or fake, and highlighting whether its decisions rely more on voiced or unvoiced regions. Building on a p…

Cited by 1SourcePDFScholar
2025

Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake Detection

ICASSP 2025accepted

In speech deepfake detection, one of the critical aspects is developing detectors able to generalize on unseen data and distinguish fake signals across different datasets. Common approaches to this challenge involve incorporating diverse data into the training process or fine-tuning models on unseen…

Cited by 0SourceScholar
2025

Leveraging Mixture of Experts for Improved Speech Deepfake Detection

ICASSP 2025accepted

Speech deepfakes pose a significant threat to personal security and content authenticity. Several detectors have been proposed in the literature, and one of the primary challenges these systems have to face is the generalization over unseen data to identify fake signals across a wide range of datase…

Cited by 0SourceScholar
2024

A One-Class Approach to Detect Super-Resolution Satellite Imagery with Spectral Features

ICASSP 2024accepted

Satellite imagery has a vital role in many applications and several techniques exist to enhance its quality. An example is given by image Super Resolution (SR), which aims at increasing the pixel resolution to recover lost high-frequency details. Due to their importance, satellite images are also a…

Cited by 0SourceScholar
2024

Mdrt: Multi-Domain Synthetic Speech Localization

ICASSP 2024accepted

With recent advancements in generating synthetic speech, tools to generate high-quality synthetic speech impersonating any human speaker are easily available. Several incidents report misuse of high-quality synthetic speech for spreading misinformation and for large-scale financial frauds. Many meth…

Cited by 0SourceScholar
2024

Water Leak Detection via Domain Adaptation

ICASSP 2024accepted

Outdated infrastructure contributes to significant water wastage, where leaks can represent as much as 30% of urban water supply losses. Rapid and precise leak detection is therefore crucial for economic and environmental reasons. Data-driven methods have emerged as promising solutions to detect wat…

Cited by 0SourceScholar
2023

ASSD: Synthetic Speech Detection in the AAC Compressed Domain

ICASSP 2023accepted

Synthetic human speech signals have become very easy to generate given modern text-to-speech methods. When these signals are shared on social media they are often compressed using the Advanced Audio Coding (AAC) standard. Our goal is to study if a small set of coding metadata contained in the AAC co…

Cited by 0SourceScholar
2023

Water Leak Detection and Localization Using Convolutional Autoencoders

ICASSP 2023accepted

Water is a valuable resource that has to be handled appropriately. However, a significant volume of water is wasted annually due to leaks in Water Distribution Networks (WDNs). This emphasizes the necessity for reliable and effective leak detection and localization systems. Several types of solution…

Cited by 0SourceScholar
2022

A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech Recording

ICASSP 2022accepted

Speech audio acquisitions exhibit different quality and reverberation properties depending on the recording setup and environment. For this reason, it is expected that speech analysis systems that work correctly on certain audio recordings may fail on others acquired in different acoustic contexts.…

Cited by 0SourceScholar
2022

Deepfake Speech Detection Through Emotion Recognition: A Semantic Approach

ICASSP 2022accepted

In recent years, audio and video deepfake technology has advanced relentlessly, severely impacting people’s reputation and reliability. Several factors have facilitated the growing deepfake threat. On the one hand, the hyper-connected society of social and mass media enables the spread of multimedia…

Cited by 0SourceScholar
2022

Forensic Analysis and Localization of Multiply Compressed MP3 Audio Using Transformers

ICASSP 2022accepted

Audio signals are often stored and transmitted in compressed formats. Among the many available audio compression schemes, MPEG-1 Audio Layer III (MP3) is very popular and widely used. Since MP3 is lossy it leaves characteristic traces in the compressed audio which can be used forensically to expose…

Cited by 0SourceScholar
2022

Panchromatic Imagery Copy-Paste Localization Through Data-Driven Sensor Attribution

ICASSP 2022accepted

Overhead images can be obtained using different acquisition and processing techniques, and they are becoming more and more popular. As with common photographs, they can be forged and manipulated by malicious users. However, not all image forensics methods tailored to normal photos can be successfull…

Cited by 0SourceScholar
2019

"Hello? Who Am I Talking to?" A Shallow CNN Approach for Human vs. Bot Speech Classification

ICASSP 2019accepted

Automatic speech generation algorithms, enhanced by deep learning techniques, enable an increasingly seamless and immediate machine-to-human interaction. As a result, the latest generation of phone-calling bots sounds more convincingly human than previous generations. The application of this technol…

Cited by 0SourceScholar
2019

Shadow Removal Detection and Localization for Forensics Analysis

ICASSP 2019accepted

The recent advancements in image processing and computer vision allow realistic photo manipulations. In order to avoid the distribution of fake imagery, the image forensics community is working towards the development of image authenticity verification tools. Methods based on shadow analysis are par…

Cited by 0SourceScholar
2018

Estimation of the Sound Field at Arbitrary Positions in Distributed Microphone Networks Based on Distributed Ray Space Transform

ICASSP 2018accepted

In this paper we propose a parametric sound field reconstruction approach. In particular, the technique is based on the estimation of three parameters for each acoustic source (source position, radiation pattern and source signal) given the signals acquired by few arbitrarily placed microphone array…

Cited by 0SourceScholar
2018

Multiple Jpeg Compression Detection Through Task-Driven Non-Negative Matrix Factorization

ICASSP 2018accepted

Due to the increasingly unbridled practice of sharing visual content on the web, tracing back past history of uploaded images is getting far from being an easy task. Nonetheless, forensic analysts might be interested in probing digital history of content published on the web to assess its authentici…

Cited by 0SourceScholar
2016

A linear operator for the computation of soundfield maps

ICASSP 2016accepted

In the process of soundfield imaging, as defined in the literature, a microphone array is subdivided into overlapping sub-arrays and soundfield images are obtained by juxtaposition of spatial spectra computed from individual subarray data. In this paper we show that the whole process can be convenie…

Cited by 0SourceScholar
2016

A low-cost solution to 3D pinna modeling for HRTF prediction

ICASSP 2016accepted

We propose an infrared (IR) stereo-vision system for estimating the 3D model of the pinna, based on low-cost devices. A commercial IR calibrated stereo camera is used in conjunction with a structured IR light projector, to acquire highly textured snapshots of the pinna. A point cloud is computed for…

Cited by 0SourceScholar
2016

Fast keypoint detection in video sequences

ICASSP 2016accepted

Several computer vision tasks exploit a succinct representation of the visual content in the form of sets of local features. Given an input image, feature extraction algorithms identify keypoints and assign to each of them a descriptor, based on the characteristics of the surrounding visual content.…

Cited by 0SourceScholar
2016

Phylogenetic analysis of near-duplicate images using processing age metrics

ICASSP 2016accepted

Recent researches on image forensics have led to the design of algorithms to study the phylogenetic relationship between near-duplicate (ND) images. The proposed solutions aim at reconstructing the image phylogeny tree (IPT), and they have immediate applications in security, law and copyright enforc…

Cited by 15SourceScholar
2015

A Dimensional Contextual Semantic Model for music description and retrieval

ICASSP 2015accepted

Several paradigms for high-level music descriptions have been proposed to develop effective system for browsing and retrieving musical content in large repositories. Such paradigms are based on either categorical or dimensional models. The interest in dimensional models has recently grown a great de…

Cited by 0SourceScholar