← Search

Daniel Stoller

7 accepted papers

2024

LLark: A Multimodal Instruction-Following Language Model for Music

ICML 2024poster

Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio. We present LLark, an instruction-tuned multimodal model for *music* understanding. We detail our process for da…

2023

Contrastive Learning-Based Audio to Lyrics Alignment for Multiple Languages

ICASSP 2023accepted

Lyrics alignment gained considerable attention in recent years. State-of-the-art systems either re-use established speech recognition toolkits, or design end-to-end solutions involving a Connectionist Temporal Classification (CTC) loss. However, both approaches suffer from specific weaknesses: toolk…

Cited by 0SourceScholar
2020

Seq-U-Net: A One-Dimensional Causal U-Net for Efficient Sequence Modelling

IJCAI 2020poster

Convolutional neural networks (CNNs) with dilated filters such as the Wavenet or the Temporal Convolutional Network (TCN) have shown good results in a variety of sequence modelling tasks. While their receptive field grows exponentially with the number of layers, computing the convolutions over very…

2020

Training Generative Adversarial Networks from Incomplete Observations using Factorised Discriminators

ICLR 2020poster

Generative adversarial networks (GANs) have shown great success in applications such as image generation and inpainting. However, they typically require large datasets, which are often not available, especially in the context of prediction tasks such as image segmentation that require labels. Theref…

Cited by 2SourcecodeScholar
2019

End-to-end Lyrics Alignment for Polyphonic Music Using an Audio-to-character Recognition Model

ICASSP 2019accepted

Time-aligned lyrics can enrich the music listening experience by enabling karaoke, text-based song retrieval and intra-song navigation, and other applications. Compared to text-to-speech alignment, lyrics alignment remains highly challenging, despite many attempts to combine numerous sub-modules inc…

Cited by 0SourceScholar
2018

Adversarial Semi-Supervised Audio Source Separation Applied to Singing Voice Extraction

ICASSP 2018accepted

The state of the art in music source separation employs neural networks trained in a supervised fashion on multi-track databases to estimate the sources from a given mixture. With only few datasets available, often extensive data augmentation is used to combat overfitting. Mixing random tracks, howe…

Cited by 0SourceScholar