← Search

Massimiliano Todisco

15 accepted papers

2024

Speaker Anonymization Using Neural Audio Codec Language Models

ICASSP 2024accepted

The vast majority of approaches to speaker anonymization involve the extraction of fundamental frequency estimates, linguistic features and a speaker embedding which is perturbed to obfuscate the speaker identity before an anonymized speech waveform is resynthesized using a vocoder. Recent work has…

Cited by 0SourceScholar
2024

Spoofing Attack Augmentation: Can Differently-Trained Attack Models Improve Generalisation?

ICASSP 2024accepted

A reliable deepfake detector or spoofing countermeasure (CM) should be robust in the face of unpredictable spoofing attacks. To encourage the learning of more generaliseable artefacts, rather than those specific only to known attacks, CMs are usually exposed to a broad variety of different attacks d…

Cited by 0SourceScholar
2024

Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset

ICASSP 2024accepted

The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech coll…

Cited by 0SourceScholar
2023

Can Spoofing Countermeasure And Speaker Verification Systems Be Jointly Optimised?

ICASSP 2023accepted

Spoofing countermeasure (CM) and automatic speaker verification (ASV) sub-systems can be used in tandem with a backend classifier as a solution to the spoofing aware speaker verification (SASV) task. The two sub-systems are typically trained independently to solve different tasks. While our previous…

Cited by 0SourceScholar
2023

StressID: a Multimodal Dataset for Stress Identification

NeurIPS 2023poster

StressID is a new dataset specifically designed for stress identification from unimodal and multimodal data. It contains videos of facial expressions, audio recordings, and physiological signals. The video and audio recordings are acquired using an RGB camera with an integrated microphone. The physi…

2022

Explaining Deep Learning Models for Spoofing and Deepfake Detection with Shapley Additive Explanations

ICASSP 2022accepted

Substantial progress in spoofing and deepfake detection has been made in recent years. Nonetheless, the community has yet to make notable inroads in providing an explanation for how a classifier produces its output. The dominance of black box spoofing detection solutions is at further odds with the…

Cited by 0SourceScholar
2022

Exploring Auditory Acoustic Features for The Diagnosis of Covid-19

ICASSP 2022accepted

The current outbreak of a coronavirus, has quickly escalated to become a serious global problem that has now been declared a Public Health Emergency of International Concern by the World Health Organization. Infectious diseases know no borders, so when it comes to controlling outbreaks, timing is ab…

Cited by 0SourceScholar
2022

Rawboost: A Raw Data Boosting and Augmentation Method Applied to Automatic Speaker Verification Anti-Spoofing

ICASSP 2022accepted

This paper introduces RawBoost, a data boosting and augmentation method for the design of more reliable spoofing detection solutions which operate directly upon raw waveform inputs. While RawBoost requires no additional data sources, e.g. noise recordings or impulse responses and is data, applicatio…

Cited by 0SourceScholar
2021

End-to-End anti-spoofing with RawNet2

ICASSP 2021accepted

Spoofing countermeasures aim to protect automatic speaker verification systems from being manipulated by spoofed speech signals. While results from the most recent ASVspoof 2019 evaluation show great potential to detect most forms of attack, some continue to evade detection. This paper reports the f…

Cited by 0SourceScholar
2020

Artificial Bandwidth Extension Using Conditional Variational Auto-encoders and Adversarial Learning

ICASSP 2020accepted

Artificial bandwidth extension (ABE) algorithms have been developed to estimate missing highband frequency components (4-8kHz) to improve quality of narrowband (0-4kHz) telephone calls. Most ABE solutions employ deep neural networks (DNNs) due to their well-known ability to model highly complex, non…

Cited by 0SourceScholar
2019

Latent Representation Learning for Artificial Bandwidth Extension Using a Conditional Variational Auto-encoder

ICASSP 2019accepted

Artificial bandwidth extension (ABE) algorithms can improve speech quality when wideband devices are used with narrowband devices or infrastructure. Most ABE solutions employ some form of memory, implying high-dimensional feature representations that increase both latency and complexity. Dimensional…

Cited by 0SourceScholar
2018

Efficient Super-Wide Bandwidth Extension Using Linear Prediction Based Analysis-Synthesis

ICASSP 2018accepted

Many smart devices now support high-quality speech communication services at super-wide bandwidths. Often, however, speech quality is degraded when they are used with networks or devices which lack super-wideband support. Artificial bandwidth extension can then be used to improve speech quality. Whi…

Cited by 0SourceScholar
2018

Exploiting Explicit Memory Inclusion for Artificial Bandwidth Extension

ICASSP 2018accepted

Artificial bandwidth extension (ABE) algorithms have been developed to improve speech quality when wideband devices are used in conjunction with narrowband devices or infrastructure. While past work points to the benefit of using contextual information or memory for ABE, an understanding of the rela…

Cited by 0SourceScholar
2017

Artificial bandwidth extension using the constant Q transform

ICASSP 2017accepted

Most artificial bandwidth extension (ABE) algorithms are based on the classical source-filter model of speech production. This approach generally requires the dual extension of each component through independent processing. Alternative approaches reported recently operate on the spectrum. With human…

Cited by 0SourceScholar
2017

RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification research

ICASSP 2017accepted

This paper describes a new database for the assessment of automatic speaker verification (ASV) vulnerabilities to spoofing attacks. In contrast to other recent data collection efforts, the new database has been designed to support the development of replay spoofing countermeasures tailored towards t…

Cited by 0SourceScholar