← Search

Antonio Almudévar

7 accepted papers

2026

SPARSE AUTOENCODERS MAKE AUDIO FOUNDATION MODELS MORE EXPLAINABLE

ICASSP 2026poster

Audio pretrained models are widely employed to solve various tasks in speech processing, sound event detection, or music information retrieval. However, the representations learned by these models are unclear, and their analysis mainly restricts to linear probing of the hidden representations. In th…

Cited by 0SourcePDFScholar
2026

There Was Never a Bottleneck in Concept Bottleneck Models

ICLR 2026poster

Deep learning representations are often difficult to interpret, which can hinder their deployment in sensitive applications. Concept Bottleneck Models (CBMs) have emerged as a promising approach to mitigate this issue by learning representations that support target task performance while ensuring th…

Cited by 0SourceScholar
2025

Aligning Multimodal Representations through an Information Bottleneck

ICML 2025poster

Contrastive losses have been extensively used as a tool for multimodal representation learning. However, it has been empirically observed that their use is not effective to learn an aligned representation space. In this paper, we argue that this phenomenon is caused by the presence of modality-spec…

Cited by 0SourcePDFScholar
2024

An Explainable Proxy Model for Multilabel Audio Segmentation

ICASSP 2024accepted

Audio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, explainable AI is a vital process for transparency of decision-making with machine learning. In this paper, we propose an e…

Cited by 0SourceScholar
2024

Unsupervised multiple domain translation through controlled Disentanglement in variational autoencoder

ICASSP 2024accepted

Unsupervised Multiple Domain Translation is the task of transforming data from one domain to other domains without having paired data to train the systems. Typically, methods based on Generative Adversarial Networks (GANs) are used to address this task. However, our proposal exclusively relies on a…

Cited by 0SourceScholar