← Search

Themos Stafylakis

15 accepted papers

2026

HYBRID PRUNING: IN-SITU COMPRESSION OF SELF-SUPERVISED SPEECH MODELS FOR SPEAKER VERIFICATION AND ANTI-SPOOFING

ICASSP 2026oral

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on resource-constrained devices. While structured pruning is a key technique for model compression, existing methods typica…

Cited by 0SourcePDFScholar
2026

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

AAAI 2026technical

Audio comprehension—including speech, non-speech sounds, and music—is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challen

Cited by 0SourcePDFScholar
2025

CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification

ICASSP 2025accepted

Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-hea…

Cited by 0SourceScholar
2024

Comparing Data Augmentation Methods for End-to-End Task-Oriented Dialog Systems

ACL 2024findings

Creating effective and reliable task-oriented dialog systems (ToDSs) is challenging, not only because of the complex structure of these systems, but also due to the scarcity of training data, especially when several modules need to be trained separately, each one with its own input/output training e…

Cited by 1SourcePDFScholar
2023

A Simple Baseline for Knowledge-Based Visual Question Answering

EMNLP 2023short main

This paper is on the problem of Knowledge-Based Visual Question Answering (KB-VQA). Recent works have emphasized the significance of incorporating both explicit (through external databases) and implicit (through LLMs) knowledge to answer questions requiring external knowledge effectively. A common l…

Cited by 0SourcecodeScholar
2023

Parameter-Efficient Transfer Learning of Pre-Trained Transformer Models for Speaker Verification Using Adapters

ICASSP 2023accepted

Recently, the pre-trained Transformer models have received a rising interest in the field of speech processing thanks to their great success in various downstream tasks. However, most fine-tuning approaches update all the parameters of the pre-trained model, which becomes prohibitive as the model si…

Cited by 0SourceScholar
2023

Speech-Based Emotion Recognition with Self-Supervised Models Using Attentive Channel-Wise Correlations and Label Smoothing

ICASSP 2023accepted

When recognizing emotions from speech, we encounter two common problems: how to optimally capture emotion-relevant information from the speech signal and how to best quantify or categorize the noisy subjective emotion labels. Self-supervised pre-trained representations can robustly capture informati…

Cited by 30SourceScholar
2020

End-to-End Architectures for ASR-Free Spoken Language Understanding

ICASSP 2020accepted

Spoken Language Understanding (SLU) is the problem of extracting the meaning from speech utterances. It is typically addressed as a two-step problem, where an Automatic Speech Recognition (ASR) model is employed to convert speech into text, followed by a Natural Language Understanding (NLU) model to…

Cited by 0SourceScholar
2019

How to Improve Your Speaker Embeddings Extractor in Generic Toolkits

ICASSP 2019accepted

Recently, speaker embeddings extracted with deep neural networks became the state-of-the-art method for speaker verification. In this paper we aim to facilitate its implementation on a more generic toolkit than Kaldi, which we anticipate to enable further improvements on the method. We examine sever…

Cited by 51SourceScholar
2019

Speaker Verification Using End-to-end Adversarial Language Adaptation

ICASSP 2019accepted

In this paper we investigate the use of adversarial domain adaptation for addressing the problem of language mismatch between speaker recognition corpora. In the context of speaker verification, adversarial domain adaptation methods aim at minimizing certain divergences between the distribution that…

Cited by 60SourceScholar
2018

End-to-End Audiovisual Speech Recognition

ICASSP 2018accepted

Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition. However, research on end-to-end audiovisual models is very limited. In this work, we present an end-to-end aud…

Cited by 0SourceScholar
2016

Towards PLDA-RBM based speaker recognition in mobile environment: Designing stacked/deep PLDA-RBM systems

ICASSP 2016accepted

The vast majority of text-independent speaker recognition systems rely on intermediate-sized vectors (i-vectors), which are compared by probabilistic linear discriminant analysis (PLDA). This paper proposes a PLDA-alike approach with restricted Boltzmann machines for i-vector based speaker recogniti…

Cited by 0SourceScholar
2015

JFA modeling with left-to-right structure and a new backend for text-dependent speaker recognition

ICASSP 2015accepted

This paper introduces a new formulation of Joint Factor Analysis (JFA) for text-dependent speaker recognition based on left-to-right modeling with tied mixture HMMs. It accommodates many different ways of extracting multiple features to characterize speakers (features may or may not be HMM state-dep…

Cited by 0SourceScholar