← Search

Santiago Pascual

15 accepted papers

2025

Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration With Improved Intelligibility

ICASSP 2025accepted

Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains high quality but, as we show, intelligibility can be substa…

Cited by 0SourceScholar
2024

GASS: Generalizing Audio Source Separation with Large-Scale Data

ICASSP 2024accepted

Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the potential of universal source separation is limited because most existing works focus on mixes with predominantly sound even…

Cited by 0SourceScholar
2024

Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity

ECCV 2024poster

"Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visual actions that are aligned with them, otherwise unnatural synchronization artifacts arise. Recent works have explored t…

2024

V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models

AAAI 2024technical

Building artificial intelligence (AI) systems on top of a set of foundation models (FMs) is becoming a new paradigm in AI research. Their representative and generative abilities learnt from vast amounts of data can be easily adapted and transferred to a wide range of downstream tasks without extra t…

2023

Adversarial Permutation Invariant Training for Universal Sound Separation

ICASSP 2023accepted

Universal sound separation consists of separating mixes with arbitrary sounds of different types, and permutation invariant training (PIT) is used to train source agnostic models that do so. In this work, we complement PIT with adversarial losses but find it challenging with the standard formulation…

Cited by 0SourceScholar
2023

Full-Band General Audio Synthesis with Score-Based Diffusion

ICASSP 2023accepted

Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds. Such models operate on band-limited signals and, as a result of an autoregressive approach, they are typically co…

Cited by 0SourceScholar
2022

On Loss Functions and Evaluation Metrics for Music Source Separation

ICASSP 2022accepted

We investigate which loss functions provide better separations via benchmarking an extensive set of those for music source separation. To that end, we first survey the most representative audio source separation losses we identified, to later consistently benchmark them in a controlled experimental…

Cited by 0SourceScholar
2021

Automatic Multitrack Mixing With A Differentiable Mixing Console Of Neural Audio Effects

ICASSP 2021accepted

Applications of deep learning to automatic multitrack mixing are largely unexplored. This is partly due to the limited available data, coupled with the fact that such data is relatively unstructured and variable. To address these challenges, we propose a domain-inspired model with a strong inductive…

Cited by 0SourceScholar
2020

Multi-Task Self-Supervised Learning for Robust Speech Recognition

ICASSP 2020accepted

Despite the growing interest in unsupervised learning, extracting meaningful knowledge from unlabelled audio remains an open challenge. To take a step in this direction, we recently proposed a problem-agnostic speech encoder (PASE), that combines a convolutional encoder followed by multiple neural n…

Cited by 0SourceScholar
2019

Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion

NeurIPS 2019poster

End-to-end models for raw audio generation are a challenge, specially if they have to work with non-parallel data, which is a desirable setup in many situations. Voice conversion, in which a model has to impersonate a speaker in a recording, is one of those situations. In this paper, we propose Blow…

2019

Wav2Pix: Speech-conditioned Face Generation Using Generative Adversarial Networks

ICASSP 2019accepted

Speech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial Network (GAN) with raw speech input. We propose a deep neura…

Cited by 0SourceScholar
2018

Language and Noise Transfer in Speech Enhancement Generative Adversarial Network

ICASSP 2018accepted

Speech enhancement deep learning systems usually require large amounts of training data to operate in broad conditions or real applications. This makes the adaptability of those systems into new, low resource environments an important topic. In this work, we present the results of adapting a speech…

Cited by 0SourceScholar