← Search

Mark B. Sandler

20 accepted papers

2023

Rigid-Body Sound Synthesis with Differentiable Modal Resonators

ICASSP 2023accepted

Physical models of rigid bodies are used for sound synthesis in applications from virtual environments to music production. Traditional methods, such as modal synthesis, often rely on computationally expensive numerical solvers, while recent deep learning approaches are limited by post-processing of…

Cited by 14SourceScholar
2022

Violinist Identification Using Note-Level Timbre Feature Distributions

ICASSP 2022accepted

Modelling musical performers’ individual playing styles based on audio features is important for music education, music expression analysis and music generation. In violin performance, the perception of playing styles are mainly affected by the characteristic musical timbre, which is mostly determin…

Cited by 0SourceScholar
2020

A Study on the Transferability of Adversarial Attacks in Sound Event Classification

ICASSP 2020accepted

An adversarial attack is an algorithm that perturbs the input of a machine learning model in an intelligent way in order to change the output of the model. An important property of adversarial attacks is transferability. According to this property, it is possible to generate adversarial perturbation…

Cited by 15SourceScholar
2020

The Fifthnet Chroma Extractor

ICASSP 2020accepted

Deep Learning (DL) is commonly used in music processing tasks such as Automatic Chord Recognition (ACR), for which Convolutional Neural Networks (CNNs) are popular tools. Compression of CNNs has become a research topic of interest, focused on post-pruning of learnt networks and development of less e…

Cited by 0SourceScholar
2019

Comparing CQT and Reassignment Based Chroma Features for Template-based Automatic Chord Recognition

ICASSP 2019accepted

Automatic Chord Recognition (ACR) seeks to extract chords from musical signals. Recently, deep neural network (DNN) approaches have become popular for this task, being employed for feature extraction and sequence modelling. Traditionally, the most important steps in ACR were extraction of chroma fea…

Cited by 0SourceScholar
2018

Shift-Invariant Kernel Additive Modelling for Audio Source Separation

ICASSP 2018accepted

A major goal in blind source separation to identify and separate sources is to model their inherent characteristics. While most state-of-the-art approaches are supervised methods trained on large datasets, interest in non-data-driven approaches such as Kernel Additive Modelling (KAM) remains high du…

Cited by 0SourceScholar
2018

Similarity Measures for Vocal-Based Drum Sample Retrieval Using Deep Convolutional Auto-Encoders

ICASSP 2018accepted

The expressive nature of the voice provides a powerful medium for communicating sonic ideas, motivating recent research on methods for query by vocalisation. Meanwhile, deep learning methods have demonstrated state-of-the-art results for matching vocal imitations to imitated sounds, yet little is kn…

Cited by 0SourceScholar
2017

Convolutional recurrent neural networks for music classification

ICASSP 2017accepted

We introduce a convolutional recurrent neural network (CRNN) for music tagging. CRNNs take advantage of convolutional neural networks (CNNs) for local feature extraction and recurrent neural networks for temporal summarisation of the extracted features. We compare CRNN with three CNN structures that…

Cited by 0SourceScholar
2017

Improved template based chord recognition using the CRP feature

ICASSP 2017accepted

The task of chord recognition in music signals is often based upon pattern matching in chromagrams. Many variants of chroma exist and quality of chord recognition is related to the feature employed. Chroma Reduced Pitch (CRP) features are interesting in this context as they were designed to improve…

Cited by 0SourceScholar
2017

Interference reduction in music recordings combining Kernel Additive Modelling and Non-Negative Matrix Factorization

ICASSP 2017accepted

In live and studio recordings unexpected sound events often lead to interferences in the signal. For non-stationary interferences, sound source separation techniques can be used to reduce the interference level in the recording. In this context, we present a novel approach combining the strengths of…

Cited by 0SourceScholar
2017

Structured dropout for weak label and multi-instance learning and its application to score-informed source separation

ICASSP 2017accepted

Many success stories involving deep neural networks are instances of supervised learning, where available labels power gradient-based learning methods. Creating such labels, however, can be expensive and thus there is increasing interest in weak labels which only provide coarse information, with unc…

Cited by 0SourceScholar
2016

A score-informed shift-invariant extension of complex matrix factorization for improving the separation of overlapped partials in music recordings

ICASSP 2016accepted

Similar to non-negative matrix factorization (NMF), complex matrix factorization (CMF) can be used to decompose a given music recording into individual sound sources. In contrast to NMF, CMF models both the magnitude and phase of a source, which can improve the separation of overlapped partials. How…

Cited by 0SourceScholar
2016

Estimation of the reliability of multiple rhythm features extraction from a single descriptor

ICASSP 2016accepted

The design of systems for automatic audio feature extraction is a central aspect of the field of Music Information Retrieval. However, feature extraction systems often do not provide an indication of the reliability of the corresponding feature. Nevertheless, the provision of a reliability or confid…

Cited by 3SourceScholar
2015

A dynamic programming variant of non-negative matrix deconvolution for the transcription of struck string instruments

ICASSP 2015accepted

Given a musical audio recording, the goal of music transcription is to determine a score-like representation of the piece underlying the recording. Most current transcription methods employ variants of non-negative matrix factorization (NMF), which often fails to robustly model instruments producing…

Cited by 0SourceScholar
2015

Non-negative matrix factorisation incorporating greedy hellinger sparse coding applied to polyphonic music transcription

ICASSP 2015accepted

Non-negative Matrix Factorisation (NMF) is a commonly used tool in many musical signal processing tasks, including Automatic Music Transcription (AMT). However unsupervised NMF is seen to be problematic in this context, and harmonically constrained variants of NMF have been proposed. While useful, t…

Cited by 6SourceScholar
2015

On the use of the tempogram to describe audio content and its application to Music structural segmentation

ICASSP 2015accepted

This paper presents a new set of audio features to describe music content based on tempo cues. Tempogram, a mid-level representation of tempo information, is constructed to characterize tempo variation and local pulse in the audio signal. We introduce a collection of novel tempogram-based features i…

Cited by 0SourceScholar