← Search

Xavier Alameda-Pineda

28 accepted papers

2026

MODELING STRATEGIES FOR SPEECH ENHANCEMENT IN THE LATENT SPACE OF A NEURAL AUDIO CODEC

ICASSP 2026poster

Neural audio codecs (NACs) provide compact latent speech representations in the form of sequences of continuous vectors or discrete tokens. In this work, we investigate how these two types of speech representations compare when used as training targets for supervised speech enhancement. We consider…

Cited by 0SourcePDFScholar
2025

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

ICASSP 2025accepted

This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key attributes, such as speaker identity, pitch, content, loudness, signal-to-noise rat…

Cited by 0SourceScholar
2025

Diffusion-based Unsupervised Audio-visual Speech Enhancement

ICASSP 2025accepted

This paper proposes a new unsupervised audiovisual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion model is pre-trained on clean speech conditioned on correspondin…

Cited by 12SourceScholar
2025

MEGA: Masked Generative Autoencoder for Human Mesh Recovery

CVPR 2025poster

Human Mesh Recovery (HMR) from a single RGB image is a highly ambiguous problem, as an infinite set of 3D interpretations can explain the 2D observation equally well. Nevertheless, most HMR methods overlook this issue and make a single prediction without accounting for this ambiguity. A few approach…

Cited by 1SourcePDFScholar
2024

A Weighted-Variance Variational Autoencoder Model for Speech Enhancement

ICASSP 2024accepted

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed for the generative model, where the speech information is encoded in the varianc…

Cited by 0SourceScholar
2024

Lost and Found: Overcoming Detector Failures in Online Multi-Object Tracking

ECCV 2024poster

"Multi-object tracking (MOT) endeavors to precisely estimate the positions and identities of multiple objects over time. The prevailing approach, tracking-by-detection (TbD), first detects objects and then links detections, resulting in a simple yet effective method. However, contemporary detectors…

2024

VQ-HPS: Human Pose and Shape Estimation in a Vector-Quantized Latent Space

ECCV 2024poster

"Previous works on Human Pose and Shape Estimation (HPSE) from RGB images can be broadly categorized into two main groups: parametric and non-parametric approaches. Parametric techniques leverage a low-dimensional statistical body model for realistic results, whereas recent non-parametric methods ac…

2023

Semi-Supervised Learning Made Simple With Self-Supervised Clustering

CVPR 2023poster

Self-supervised learning models have been shown to learn rich visual representations without requiring human annotations. However, in many real-world scenarios, labels are partially available, motivating a recent line of work on semi-supervised methods inspired by self-supervised principles. In this…

2023

Speech Modeling with a Hierarchical Transformer Dynamical VAE

ICASSP 2023accepted

The dynamical variational autoencoders (DVAEs) are a family of latent-variable deep generative models that extends the VAE to model a sequence of observed data and a corresponding sequence of latent vectors. In almost all the DVAEs of the literature, the temporal dependencies within each sequence an…

Cited by 0SourceScholar
2022

A Proposal-Based Paradigm for Self-Supervised Sound Source Localization in Videos

CVPR 2022poster

Humans can easily recognize where and how the sound is produced via watching a scene and listening to corresponding audio cues. To achieve such cross-modal perception on machines, existing methods only use the maps generated by interpolation operations to localize the sound source. As semantic objec…

Cited by 22PDFScholar
2022

Active Contrastive Set Mining for Robust Audio-Visual Instance Discrimination

IJCAI 2022poster

The recent success of audio-visual representation learning can be largely attributed to their pervasive property of audio-visual synchronization, which can be used as self-annotated supervision. As a state-of-the-art solution, Audio-Visual Instance Discrimination (AVID) extends instance discriminati…

Cited by 1SourcePDFScholar
2022

Self-Supervised Models Are Continual Learners

CVPR 2022poster

Self-supervised models have been shown to produce comparable or better visual representations than their supervised counterparts when trained offline on unlabeled data at scale. However, their efficacy is catastrophically reduced in a Continual Learning (CL) scenario where data is presented to the m…

Cited by 217PDFcodeScholar
2022

The Impact of Removing Head Movements on Audio-Visual Speech Enhancement

ICASSP 2022accepted

This paper investigates the impact of head movements on audio-visual speech enhancement (AVSE). Although being a common conversational feature, head movements have been ignored by past and recent studies: they challenge today’s learning-based methods as they often degrade the performance of models t…

Cited by 0SourceScholar
2021

Switching Variational Auto-Encoders for Noise-Agnostic Audio-Visual Speech Enhancement

ICASSP 2021accepted

Recently, audio-visual speech enhancement has been tackled in the unsupervised settings based on variational auto-encoders (VAEs), where during training only clean data is used to train a generative model for speech, which at test time is combined with a noise model, e.g. nonnegative matrix factoriz…

Cited by 0SourceScholar
2020

A Recurrent Variational Autoencoder for Speech Enhancement

ICASSP 2020accepted

This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a nonnegative matrix factorization noise model for speech enhancement. We propose…

Cited by 0SourceScholar
2020

How to Train Your Deep Multi-Object Tracker

CVPR 2020poster

The recent trend in vision-based multi-object tracking (MOT) is heading towards leveraging the representational power of deep learning to jointly learn to detect and track objects. However, existing methods train only certain sub-modules using loss functions that often do not correlate with establis…

Cited by 274PDFcodeScholar
2020

Robust Unsupervised Audio-Visual Speech Enhancement Using a Mixture of Variational Autoencoders

ICASSP 2020accepted

Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised speech enhancement. When visual data is clean, speech enhancement with audio-vi…

Cited by 0SourceScholar
2018

Accounting for Room Acoustics in Audio-Visual Multi-Speaker Tracking

ICASSP 2018accepted

Multiple-speaker tracking is a crucial task for many applications. In real-world scenarios, exploiting the complementarity between auditory and visual data enables to track people outside the visual field of view. However, practical methods must be robust to changes in acoustic conditions, e.g. reve…

Cited by 0SourceScholar
2018

DeepGUM: Learning Deep Robust Regression with a Gaussian-Uniform Mixture Model

ECCV 2018poster

In this paper we address the problem of how to robustly train a ConvNet for regression, or deep robust regression. Traditionally, deep regression employ the L2 loss function, known to be sensitive to outliers, i.e. samples that either lie at an abnormal distance away from the majority of the trainin…

Cited by 37SourcePDFScholar
2018

Every Smile Is Unique: Landmark-Guided Diverse Smile Generation

CVPR 2018poster

Each smile is unique: one person surely smiles in different ways (e.g., closing/opening the eyes or mouth). Given one input image of a neutral face, can we generate multiple smile videos with distinctive characteristics? To tackle this one-to-many video generation problem, we propose a novel deep le…

Cited by 82SourcePDFScholar
2017

An EM algorithm for joint source separation and diarisation of multichannel convolutive speech mixtures

ICASSP 2017accepted

We present a probabilistic model for joint source separation and diarisation of multichannel convolutive speech mixtures. We build upon the framework of local Gaussian model (LGM) with non-negative matrix factorization (NMF). The diarisation is introduced as a temporal labeling of each source in the…

Cited by 0SourceScholar
2017

Learning Deep Structured Multi-Scale Features using Attention-Gated CRFs for Contour Prediction

NeurIPS 2017poster

Recent works have shown that exploiting multi-scale representations deeply learned via convolutional neural networks (CNN) is of tremendous importance for accurate contour detection. This paper presents a novel approach for predicting contours which advances the state of the art in two fundamental a…

2017

Tracking a varying number of people with a visually-controlled robotic head

IROS 2017poster

Multi-person tracking with a robotic platform is one of the cornerstones of human-robot interaction. Challenges arise from occlusions, appearance changes and a time-varying number of people. Furthermore, the final system is constrained by the hardware platform: low computational capacity and limited…

Cited by 25SourceScholar
2016

An inverse-gamma source variance prior with factorized parameterization for audio source separation

ICASSP 2016accepted

In this paper we present a new statistical model for the power spectral density (PSD) of an audio signal and its application to multichannel audio source separation (MASS). The source signal is modeled with the local Gaussian model (LGM) and we propose to model its variance with an inverse-Gamma dis…

Cited by 0SourceScholar
2016

Recognizing Emotions From Abstract Paintings Using Non-Linear Matrix Completion

CVPR 2016poster

Advanced computer vision and machine learning techniques tried to automatically categorize the emotions elicited by abstract paintings with limited success. Since the annotation of the emotional content is highly resource-consuming, datasets of abstract paintings are either constrained in size or pa…

Cited by 120PDFcodeScholar
2016

Self-Adaptive Matrix Completion for Heart Rate Estimation From Face Videos Under Realistic Conditions

CVPR 2016oral

Recent studies in computer vision have shown that, while practically invisible to a human observer, skin color changes due to blood flow can be captured on face videos and, surprisingly, be used to estimate the heart rate (HR). While considerable progress has been made in the last few years, still m…

Cited by 413PDFScholar