← Search

Nathanaël Perraudin

11 accepted papers

2025

Bootstrapping Language-Audio Pre-training for Music Captioning

ICASSP 2025accepted

We introduce BLAP, a model capable of generating high-quality captions for music. BLAP leverages a fine-tuned CLAP audio encoder and a pre-trained Flan-T5 large language model. To achieve effective cross-modal alignment between music and language, BLAP utilizes a Querying Transformer, allowing us to…

Cited by 0SourceScholar
2025

Coarse-to-Fine Text-to-Music Latent Diffusion

ICASSP 2025accepted

We introduce DiscoDiff, a text-to-music generative model that utilizes two latent diffusion models to produce high-fidelity 44.1kHz music hierarchically. Our approach significantly enhances audio quality through a coarse-to-fine generation strategy, leveraging residual vector quantization from the D…

Cited by 0SourceScholar
2024

Efficient and Scalable Graph Generation through Iterative Local Expansion

ICLR 2024poster

In the realm of generative models for graphs, extensive research has been conducted. However, most existing methods struggle with large graphs due to the complexity of representing the entire joint distribution across all node pairs and capturing both global and local graph structures simultaneously…

2022

SPECTRE: Spectral Conditioning Helps to Overcome the Expressivity Limits of One-shot Graph Generators

ICML 2022spotlight

We approach the graph generation problem from a spectral perspective by first generating the dominant parts of the graph Laplacian spectrum and then building a graph matching these eigenvalues and eigenvectors. Spectral conditioning allows for direct modeling of the global and local graph structure…

2022

What You See is What You Classify: Black Box Attributions

NeurIPS 2022accept

An important step towards explaining deep image classifiers lies in the identification of image regions that contribute to individual class scores in the model's output. However, doing this accurately is a difficult task due to the black-box nature of such networks. Most existing approaches find suc…

2021

Scalable Graph Networks for Particle Simulations

AAAI 2021technical

Learning system dynamics directly from observations is a promising direction in machine learning due to its potential to significantly enhance our ability to understand physical systems. However, the dynamics of many real-world systems are challenging to learn due to the presence of nonlinear potent…

2020

DeepSphere: a graph-based spherical CNN

ICLR 2020spotlight

Designing a convolution for a spherical neural network requires a delicate tradeoff between efficiency and rotation equivariance. DeepSphere, a method based on a graph representation of the discretized sphere, strikes a controllable balance between these two desiderata. This contribution is twofold.…

Cited by 116SourcecodeScholar
2019

Adversarial Generation of Time-Frequency Features with application in audio synthesis

ICML 2019oral

Time-frequency (TF) representations provide powerful and intuitive features for the analysis of time series such as audio. But still, generative modeling of audio in the TF domain is a subtle matter. Consequently, neural audio synthesis widely relies on directly modeling the waveform and previous at…

2017

Towards stationary time-vertex signal processing

ICASSP 2017accepted

Graph-based methods for signal processing have shown promise for the analysis of data exhibiting irregular structure, such as those found in social, transportation, and sensor networks. Yet, though these systems are often dynamic, state-of-the-art methods for graph signal processing ignore the time…

Cited by 0SourceScholar
2016

PCA using graph total variation

ICASSP 2016accepted

Mining useful clusters from high dimensional data has received significant attention of the signal processing and machine learning community in the recent years. Linear and non-linear dimensionality reduction has played an important role to overcome the curse of dimensionality. However, often such m…

Cited by 0SourceScholar