← Search

Alicia Lozano-Diez

4 accepted papers

2026

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

AAAI 2026technical

Audio comprehension—including speech, non-speech sounds, and music—is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challen

Cited by 0SourcePDFScholar
2025

Personalizing Keyword Spotting with Speaker Information

ICASSP 2025accepted

Keyword spotting systems often struggle to generalize to a diverse population with various accents and age groups. To address this challenge, we propose a novel approach that integrates speaker information into keyword spotting using Feature-wise Linear Modulation (FiLM), a recent method that allows…

Cited by 0SourceScholar
2023

Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural Diarization

ICASSP 2023accepted

End-to-end diarization presents an attractive alternative to standard cascaded diarization systems because a single system can handle all aspects of the task at once. Many flavors of end-to-end models have been proposed but all of them require (so far non-existing) large amounts of annotated data fo…

Cited by 0SourceScholar
2018

DNN Based Embeddings for Language Recognition

ICASSP 2018accepted

In this work, we present a language identification (LID) system based on embeddings. In our case, an embedding is a fixed-length vector (similar to i-vector) that represents the whole utterance, but unlike i-vector it is designed to contain mostly information relevant to the target task (LID). In or…

Cited by 28SourceScholar