← Search

Simon Dixon

22 accepted papers

2026

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

ICML 2026poster

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind, remaining fragmented and narrowly focused. In this paper, we bridge this critical gap by establishing a comprehensive ecosystem for Compo…

Cited by 0SourceScholar
2025

LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment

ICASSP 2025accepted

Research in music understanding has extensively explored composition-level attributes such as key, genre, and instrumentation through advanced representations, leading to cross-modal applications using large language models. However, aspects of musical performance such as stylistic expression and te…

Cited by 0SourceScholar
2024

MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models

IJCAI 2024poster

Recent advances in text-to-music generation models have opened new avenues in musical creativity. However, the task of editing these generated music remains a significant challenge. This paper introduces a novel approach to edit music generated by such models, enabling the modification of specific a…

2024

Posterior Variance-Parameterised Gaussian Dropout: Improving Disentangled Sequential Autoencoders for Zero-Shot Voice Conversion

ICASSP 2024accepted

The class of disentangled sequential auto-encoders factorises speech into time-invariant (global) and time-variant (local) representations for speaker identity and linguistic content, respectively. Many of the existing models employ this assumption to tackle zero-shot voice conversion (VC), which co…

Cited by 0SourceScholar
2024

Unsupervised Pitch-Timbre Disentanglement of Musical Instruments Using a Jacobian Disentangled Sequential Autoencoder

ICASSP 2024accepted

Disentangled representation learning seeks to align individual dimensions or separate groups of coordinates of latent factors with attributes of observed data such that perturbing certain latent factors uniquely changes particular attributes. A main challenge in unsupervised disentanglement using au…

Cited by 0SourceScholar
2023

Disentangling the Horowitz Factor: Learning Content and Style From Expressive Piano Performance

ICASSP 2023accepted

In the Western art music tradition, expressive piano performance consists of two kinds of information: the score, with pitch and timing expressed in simple musical units along with occasional expression instructions, and the performer’s interpretation of the score, involving variations in tempo, dyn…

Cited by 0SourceScholar
2022

Towards Robust Unsupervised Disentanglement of Sequential Data — A Case Study Using Music Audio

IJCAI 2022poster

Disentangled sequential autoencoders (DSAEs) represent a class of probabilistic graphical models that describes an observed sequence with dynamic latent variables and a static latent variable. The former encode information at a frame rate identical to the observation, while the latter globally gover…

2021

Structure-Aware Audio-to-Score Alignment Using Progressively Dilated Convolutional Neural Networks

ICASSP 2021accepted

The identification of structural differences between a music performance and the score is a challenging yet integral step of audio-to-score alignment, an important subtask of music information retrieval. We present a novel method to detect such differences between the score and performance for a giv…

Cited by 0SourceScholar
2020

Seq-U-Net: A One-Dimensional Causal U-Net for Efficient Sequence Modelling

IJCAI 2020poster

Convolutional neural networks (CNNs) with dilated filters such as the Wavenet or the Temporal Convolutional Network (TCN) have shown good results in a variety of sequence modelling tasks. While their receptive field grows exponentially with the number of layers, computing the convolutions over very…

2020

Training Generative Adversarial Networks from Incomplete Observations using Factorised Discriminators

ICLR 2020poster

Generative adversarial networks (GANs) have shown great success in applications such as image generation and inpainting. However, they typically require large datasets, which are often not available, especially in the context of prediction tasks such as image segmentation that require labels. Theref…

Cited by 2SourcecodeScholar
2018

Adversarial Semi-Supervised Audio Source Separation Applied to Singing Voice Extraction

ICASSP 2018accepted

The state of the art in music source separation employs neural networks trained in a supervised fashion on multi-track databases to estimate the sources from a given mixture. With only few datasets available, often extensive data augmentation is used to combat overfitting. Mixing random tracks, howe…

Cited by 0SourceScholar
2018

Similarity Measures for Vocal-Based Drum Sample Retrieval Using Deep Convolutional Auto-Encoders

ICASSP 2018accepted

The expressive nature of the voice provides a powerful medium for communicating sonic ideas, motivating recent research on methods for query by vocalisation. Meanwhile, deep learning methods have demonstrated state-of-the-art results for matching vocal imitations to imitated sounds, yet little is kn…

Cited by 0SourceScholar
2018

Towards Complete Polyphonic Music Transcription: Integrating Multi-Pitch Detection and Rhythm Quantization

ICASSP 2018accepted

Most work on automatic transcription produces “piano roll” data with no musical interpretation of the rhythm or pitches. We present a polyphonic transcription method that converts a music audio signal into a human-readable musical score, by integrating multi-pitch detection and rhythm quantization m…

Cited by 0SourceScholar
2017

Pickup position and plucking point estimation on an electric guitar

ICASSP 2017accepted

This paper describes a technique to estimate the plucking point and magnetic pickup location along the strings of an electric guitar from a recording of an isolated guitar tone. The estimated values are calculated by minimising the difference between the magnitude spectrum of the recorded tone and t…

Cited by 0SourceScholar
2017

Towards the characterization of singing styles in world music

ICASSP 2017accepted

In this paper we focus on the characterization of singing styles in world music.We develop a set of contour features capturing pitch structure and melodic embellishments.Using these features we train a binary classifier to distinguish vocal from non-vocal contours and learn a dictionary of singing s…

Cited by 0SourceScholar
2016

Estimation of the reliability of multiple rhythm features extraction from a single descriptor

ICASSP 2016accepted

The design of systems for automatic audio feature extraction is a central aspect of the field of Music Information Retrieval. However, feature extraction systems often do not provide an indication of the reliability of the corresponding feature. Nevertheless, the provision of a reliability or confid…

Cited by 3SourceScholar
2015

A hybrid recurrent neural network for music transcription

ICASSP 2015accepted

We investigate the problem of incorporating higher-level symbolic score-like information into Automatic Music Transcription (AMT) systems to improve their performance. We use recurrent neural networks (RNNs) and their variants as music language models (MLMs) and present a generative architecture for…

Cited by 0SourceScholar
2015

Compensating for asynchronies between musical voices in score-performance alignment

ICASSP 2015accepted

The goal of score-performance synchronisation is to align a given musical score to an audio recording of a performance of the same piece. A major challenge in computing such alignments is to account for musical parameters including the local tempo or playing style. To increase the overall robustness…

Cited by 0SourceScholar