← Search

Xavier Favory

5 accepted papers

2023

Pre-Training Strategies Using Contrastive Learning and Playlist Information for Music Classification and Similarity

ICASSP 2023accepted

In this work, we investigate an approach that relies on contrastive learning and music metadata as a weak source of supervision to train music representation models. Recent studies show that contrastive learning can be used with editorial metadata (e.g., artist or album name) to learn audio represen…

Cited by 0SourceScholar
2021

Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags

ICASSP 2021accepted

Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Published approaches that consider both audio and words/tags associated with audio do not employ text processing models that ar…

Cited by 0SourceScholar
2020

Neural Percussive Synthesis Parameterised by High-Level Timbral Features

ICASSP 2020accepted

We present a deep neural network-based methodology for synthesising percussive sounds with control over high-level timbral characteristics of the sounds. This approach allows for intuitive control of a synthesizer, enabling the user to shape sounds without extensive knowledge of signal processing. W…

Cited by 26SourceScholar
2019

Learning Sound Event Classifiers from Web Audio with Noisy Labels

ICASSP 2019accepted

As sound event classification moves towards larger datasets, issues of label noise become inevitable. Web sites can supply large volumes of user-contributed audio and metadata, but inferring labels from this metadata introduces errors due to unreliable inputs, and limitations in the mapping. There i…

Cited by 0SourceScholar
2015

The role of glottal source parameters for high-quality transformation of perceptual age

ICASSP 2015accepted

The intuitive control of voice transformation (e.g., age/sex, emotions) is useful to extend the expressive repertoire of a voice. This paper explores the role of glottal source parameters for the control of voice transformation. First, the SVLN speech synthesizer (Separation of the Vocal-tract with…

Cited by 2SourceScholar