← Search

Luciana Ferrer

10 accepted papers

2025

LSCD: Lomb--Scargle Conditioned Diffusion for Time series Imputation

ICML 2025poster

Time series with missing or irregularly sampled data are a persistent challenge in machine learning. Many methods operate on the frequency-domain, relying on the Fast Fourier Transform (FFT) which assumes uniform sampling, therefore requiring prior interpolation that can distort the spectra. To addr…

Cited by 0SourcePDFScholar
2023

Study on the Fairness of Speaker Verification Systems Across Accent and Gender Groups

ICASSP 2023accepted

Speaker verification (SV) systems are currently used for consequential tasks like giving access to bank accounts or making forensic decisions. Ensuring that these systems are fair and do not disfavor any particular group is crucial. In this work, we analyze the performance of two X-vector-based SV s…

Cited by 0SourceScholar
2022

A Transfer Learning Approach for Pronunciation Scoring

ICASSP 2022accepted

Phone-level pronunciation scoring is a challenging task, with performance far from that of human annotators. Standard systems generate a score for each phone in a phrase using models trained for automatic speech recognition (ASR) with native data only. Better performance has been shown when using sy…

Cited by 0SourceScholar
2022

Study of Positional Encoding Approaches for Audio Spectrogram Transformers

ICASSP 2022accepted

Transformers have revolutionized the world of deep learning, specially in the field of natural language processing. Recently, the Audio Spectrogram Transformer (AST) was proposed for audio classification, leading to state of the art results in several datasets. However, in order for ASTs to outperfo…

Cited by 0SourceScholar
2020

Fusion Approaches for Emotion Recognition from Speech Using Acoustic and Text-Based Features

ICASSP 2020accepted

In this paper, we study different approaches for classifying emotions from speech using acoustic and text-based features. We propose to obtain contextualized word embeddings with BERT to represent the information contained in speech transcriptions and show that this results in better performance tha…

Cited by 0SourceScholar
2019

Analysis and Mitigation of Vocal Effort Variations in Speaker Recognition

ICASSP 2019accepted

In this work, we assess the impact of vocal effort on discrimination and calibration performance of a state-of-the-art speaker recognition system. We analyze three levels of vocal effort (low, normal, and high) from the SRI-FRTIV corpus. We use a deep neural network (DNN) speaker embeddings system w…

Cited by 0SourceScholar
2016

Exploring the role of phonetic bottleneck features for speaker and language recognition

ICASSP 2016accepted

Using bottleneck features extracted from a deep neural network (DNN) trained to predict senone posteriors has resulted in new, state-of-the-art technology for language and speaker identification. For language identification, the features' dense phonetic information is believed to enable improved per…

Cited by 0SourceScholar