← Search

Curtis Hawthorne

7 accepted papers

2022

General-purpose, long-context autoregressive modeling with Perceiver AR

ICML 2022spotlight

Real-world data is high-dimensional: a book, image, or musical performance can easily contain hundreds of thousands of elements even after compression. However, the most commonly used autoregressive models, Transformers, are prohibitively expensive to scale to the number of inputs and layers needed…

2022

Improving Source Separation by Explicitly Modeling Dependencies between Sources

ICASSP 2022accepted

We propose a new method for training a supervised source separation system that aims to learn the interdependent relationships between all combinations of sources in a mixture. Rather than independently estimating each source from a mix, we reframe the source separation problem as an Orderless Neura…

Cited by 0SourceScholar
2022

MT3: Multi-Task Multitrack Music Transcription

ICLR 2022spotlight

Automatic Music Transcription (AMT), inferring musical notes from raw audio, is a challenging task at the core of music understanding. Unlike Automatic Speech Recognition (ASR), which typically focuses on the words of a single speaker, AMT often requires transcribing multiple instruments simultaneou…

2020

Encoding Musical Style with Transformer Autoencoders

ICML 2020poster

We consider the problem of learning high-level controls over the global structure of generated sequences, particularly in the context of symbolic music generation with complex language models. In this work, we present the Transformer autoencoder, which aggregates encodings of the input data across t…

Cited by 136SourcePDFScholar
2019

Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset

ICLR 2019oral

Generating musical audio directly with neural networks is notoriously difficult because it requires coherently modeling structure at many different timescales. Fortunately, most music is also highly structured and can be represented as discrete note events played on musical instruments. Herein, we s…

Cited by 629SourcePDFScholar
2019

Music Transformer: Generating Music with Long-Term Structure

ICLR 2019poster

Music relies heavily on repetition to build structure and meaning. Self-reference occurs on multiple timescales, from motifs to phrases to reusing of entire sections of music, such as in pieces with ABA structure. The Transformer (Vaswani et al., 2017), a sequence model based on self-attention, ha…

Cited by 0SourcePDFScholar
2018

A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music

ICML 2018oral

The Variational Autoencoder (VAE) has proven to be an effective model for producing semantically meaningful latent representations for natural data. However, it has thus far seen limited application to sequential data, and, as we demonstrate, existing recurrent VAE models have difficulty modeling se…

Cited by 676SourcePDFScholar