← Search

Tomi Kinnunen

10 accepted papers

2026

WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection

ICASSP 2026poster

Modern front-end design for speech deepfake detection relies on full fine-tuning of large pre-trained models like XLSR. However, this approach is not parameter-efficient and may lead to suboptimal generalization to realistic, in-the-wild data types. To address these limitations, we introduce a new f…

Cited by 0SourcePDFScholar
2023

Learnable Frontends That Do Not Learn: Quantifying Sensitivity To Filterbank Initialisation

ICASSP 2023accepted

While much of modern speech and audio processing relies on deep neural networks trained using fixed audio representations, recent studies suggest great potential in acoustic frontends learnt jointly with a backend. In this study, we focus specifically on learnable filterbanks. Prior studies have rep…

Cited by 0SourceScholar
2019

Can We Use Speaker Recognition Technology to Attack Itself? Enhancing Mimicry Attacks Using Automatic Target Speaker Selection

ICASSP 2019accepted

We consider technology-assisted mimicry attacks in the context of automatic speaker verification (ASV). We use ASV itself to select targeted speakers to be attacked by human-based mimicry. We recorded 6 naive mimics for whom we select target celebrities from VoxCeleb1 and VoxCeleb2 corpora (7,365 po…

Cited by 0SourceScholar
2019

Who Do I Sound like? Showcasing Speaker Recognition Technology by Youtube Voice Search

ICASSP 2019accepted

The popularization of science can often be disregarded by scientists as it may be challenging to put highly sophisticated research into words that general public can understand. This work aims to help presenting speaker recognition research to public by proposing a publicly appealing concept for sho…

Cited by 0SourceScholar
2017

Effects of gender information in text-independent and text-dependent speaker verification

ICASSP 2017accepted

It is well-known that for speaker recognition task, gender-dependent acoustic modeling performs better than gender-independent modeling. The practice is to use the gender ground-truth and to train gender-dependent models. However, such information is not necessarily available, especially if speakers…

Cited by 0SourceScholar
2017

Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformation

ICASSP 2017accepted

Text-independent speaker verification (recognizing speakers regardless of content) and non-parallel voice conversion (transforming voice identities without requiring content-matched training utterances) are related problems. We adopt i-vector method to voice conversion. An i-vector is a fixed-dimens…

Cited by 0SourceScholar
2017

RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification research

ICASSP 2017accepted

This paper describes a new database for the assessment of automatic speaker verification (ASV) vulnerabilities to spoofing attacks. In contrast to other recent data collection efforts, the new database has been designed to support the development of replay spoofing countermeasures tailored towards t…

Cited by 123SourceScholar