← Search

Mark D. Plumbley

35 accepted papers

2026

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) represent an important frontier in multimodal AI, addressing diverse audio tasks. Recently, post-training of LALMs has received increasing attention due to significant performance improvements over foundation models. While single-stage post-training such as reinfo…

Cited by 0SourcecodeScholar
2025

A decade of DCASE: Achievements, practices, evaluations and future challenges

ICASSP 2025accepted

This paper introduces briefly the history and growth of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, workshop, research area and research community. Created in 2013 as a data evaluation challenge, DCASE has become a major research topic in the Audio and Acoustic…

Cited by 0SourceScholar
2025

FlowSep: Language-Queried Sound Separation with Rectified Flow Matching

ICASSP 2025accepted

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds and minimize interference from other sources. However, these…

Cited by 0SourceScholar
2025

Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions

ICASSP 2025accepted

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work…

Cited by 0SourceScholar
2024

Learning Temporal Resolution in Spectrogram for Audio Classification

AAAI 2024technical

The audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time Fourier Transform (STFT). Previous works generally assume t…

2024

Retrieval-Augmented Text-to-Audio Generation

ICASSP 2024accepted

Despite recent progress in text-to-audio (TTA) generation, we show that the state-of-the-art models, such as AudioLDM, trained on datasets with an imbalanced class distribution, such as AudioCaps, are biased in their generation performance. Specifically, they excel in generating common audio classes…

Cited by 0SourceScholar
2023

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

ICML 2023poster

Text-to-audio (TTA) systems have recently gained attention for their ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a lat…

2023

Simple Pooling Front-Ends for Efficient Audio Classification

ICASSP 2023accepted

Recently, there has been increasing interest in building efficient audio neural networks for on-device scenarios. Most existing approaches are designed to reduce the size of audio neural networks using methods such as model pruning. In this work, we show that instead of reducing model size using com…

Cited by 0SourceScholar
2022

A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection

ICASSP 2022accepted

Polyphonic sound event localization and detection (SELD) aims at detecting types of sound events with corresponding temporal activities and spatial locations. In this paper, a trackwise ensemble event independent network with a novel data augmentation method is proposed. The proposed model is based…

Cited by 0SourceScholar
2021

An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection

ICASSP 2021accepted

Polyphonic sound event localization and detection (SELD), which jointly performs sound event detection (SED) and direction-of-arrival (DoA) estimation, detects the type and occurrence time of sound events as well as their corresponding DoA angles simultaneously. We study the SELD task from a multi-t…

Cited by 0SourceScholar
2020

Learning With Out-of-Distribution Data for Audio Classification

ICASSP 2020accepted

In supervised machine learning, the assumption that training data is labelled correctly is not always satisfied. In this paper, we investigate an instance of labelling error for classification tasks in which the dataset is corrupted with out-of-distribution (OOD) instances: data that does not belong…

Cited by 0SourceScholar
2020

Source Separation with Weakly Labelled Data: an Approach to Computational Auditory Scene Analysis

ICASSP 2020accepted

Source separation is the task of separating an audio recording into individual sound sources. Source separation is fundamental for computational auditory scene analysis. Previous work on source separation has focused on separating particular sound classes such as speech and music. Much previous work…

Cited by 0SourceScholar
2019

Acoustic Event Detection from Weakly Labeled Data Using Auditory Salience

ICASSP 2019accepted

Acoustic Event Detection (AED) is an important task of machine listening which, in recent years, has been addressed using common machine learning methods like Non-negative Matrix Factorization (NMF) or deep learning. However, most of these approaches do not take into consideration the way that human…

Cited by 0SourceScholar
2019

Acoustic Scene Generation with Conditional Samplernn

ICASSP 2019accepted

Acoustic scene generation (ASG) is a task to generate waveforms for acoustic scenes. ASG can be used to generate audio scenes for movies and computer games. Recently, neural networks such as SampleRNN have been used for speech and music generation. However, ASG is more challenging due to its wide va…

Cited by 0SourceScholar
2019

Attention-based Atrous Convolutional Neural Networks: Visualisation and Understanding Perspectives of Acoustic Scenes

ICASSP 2019accepted

The goal of Acoustic Scene Classification (ASC) is to recognise the environment in which an audio waveform has been recorded. Recently, deep neural networks have been applied to ASC and have achieved state-of-the-art performance. However, few works have investigated how to visualise and understand w…

Cited by 0SourceScholar
2019

Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and Challenge

ICASSP 2019accepted

Humans are able to identify a large number of environmental sounds and categorise them according to high-level semantic categories, e.g. urban sounds or music. They are also capable of generalising from past experience to new sounds when applying these categories. In this paper we report on the crea…

Cited by 0SourceScholar
2019

Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised Clustering

ICASSP 2019accepted

Sound event detection (SED) methods typically rely on either strongly labelled data or weakly labelled data. As an alternative, sequentially labelled data (SLD) was proposed. In SLD, the events and the order of events in audio clips are known, without knowing the occurrence time of events. This pape…

Cited by 0SourceScholar
2018

A Joint Separation-Classification Model for Sound Event Detection of Weakly Labelled Data

ICASSP 2018accepted

Source separation (SS) aims to separate individual sources from an audio recording. Sound event detection (SED) aims to detect sound events from an audio recording. We propose a joint separation-classification (JSC) model trained only on weakly labelled audio data, that is, only the tags of an audio…

Cited by 0SourceScholar
2018

Audio Set Classification with Attention Model: A Probabilistic Perspective

ICASSP 2018accepted

This paper investigates the Audio Set classification. Audio Set is a large scale weakly labelled dataset (WLD) of audio clips. In WLD only the presence of a label is known, without knowing the happening time of the labels. We propose an attention model to solve this WLD problem and explain the atten…

Cited by 0SourceScholar
2018

BSS Eval or Peass? Predicting the Perception of Singing-Voice Separation

ICASSP 2018accepted

There is some uncertainty as to whether objective metrics for predicting the perceived quality of audio source separation are sufficiently accurate. This issue was investigated by employing a revised experimental methodology to collect subjective ratings of sound quality and interference of singing-…

Cited by 0SourceScholar
2018

Large-Scale Weakly Supervised Audio Classification Using Gated Convolutional Neural Network

ICASSP 2018accepted

In this paper, we present a gated convolutional neural network and a temporal attention-based localization method for audio classification, which won the 1st place in the large-scale weakly supervised sound event detection task of Detection and Classification of Acoustic Scenes and Events (DCASE) 20…

Cited by 0SourceScholar
2018

Orthogonality-Regularized Masked NMF for Learning on Weakly Labeled Audio Data

ICASSP 2018accepted

Non-negative Matrix Factorization (NMF) is a well established tool for audio analysis. However, it is not well suited for learning on weakly labeled data, i.e. data where the exact timestamp of the sound of interest is not known. In this paper we propose a novel extension to NMF, that allows it to e…

Cited by 0SourceScholar
2018

Synthesis of Images by Two-Stage Generative Adversarial Networks

ICASSP 2018accepted

In this paper, we propose a divide-and-conquer approach using two generative adversarial networks (GANs) to explore how a machine can draw colorful pictures (bird) using a small amount of training data. In our work, we simulate the procedure of an artist drawing a picture, where one begins with draw…

Cited by 0SourceScholar
2017

A greedy algorithm with learned statistics for sparse signal reconstruction

ICASSP 2017accepted

We address the problem of sparse signal reconstruction from a few noisy samples. Recently, a Covariance-Assisted Matching Pursuit (CAMP) algorithm has been proposed, improving the sparse coefficient update step of the classic Orthogonal Matching Pursuit (OMP) algorithm. CAMP allows the a-priori mean…

Cited by 0SourceScholar
2017

A joint detection-classification model for audio tagging of weakly labelled data

ICASSP 2017accepted

Audio tagging aims to assign one or several tags to an audio clip. Most of the datasets are weakly labelled, which means only the tags of the clip are known, without knowing the occurrence time of the tags. The labeling of an audio clip is often based on the audio events in the clip and no event lev…

Cited by 0SourceScholar
2017

Assessment of musical noise using localization of isolated peaks in time-frequency domain

ICASSP 2017accepted

Musical noise is a recurrent issue that appears in spectral techniques for denoising or blind source separation. Due to localised errors of estimation, isolated peaks may appear in the processed spectrograms, resulting in annoying tonal sounds after synthesis known as “musical noise”. In this paper,…

Cited by 0SourceScholar
2017

Fast tagging of natural sounds using marginal co-regularization

ICASSP 2017accepted

Automatic and fast tagging of natural sounds in audio collections is a very challenging task due to wide acoustic variations, the large number of possible tags, the incomplete and ambiguous tags provided by different labellers. To handle these problems, we use a co-regularization approach to learn a…

Cited by 0SourceScholar
2016

Detection of overlapping acoustic events using a temporally-constrained probabilistic model

ICASSP 2016accepted

In this paper, a system for overlapping acoustic event detection is proposed, which models the temporal evolution of sound events. The system is based on probabilistic latent component analysis, supporting the use of a sound event dictionary where each exemplar consists of a succession of spectral t…

Cited by 0SourceScholar
2015

A dynamic programming variant of non-negative matrix deconvolution for the transcription of struck string instruments

ICASSP 2015accepted

Given a musical audio recording, the goal of music transcription is to determine a score-like representation of the piece underlying the recording. Most current transcription methods employ variants of non-negative matrix factorization (NMF), which often fails to robustly model instruments producing…

Cited by 0SourceScholar
2015

Non-negative matrix factorisation incorporating greedy hellinger sparse coding applied to polyphonic music transcription

ICASSP 2015accepted

Non-negative Matrix Factorisation (NMF) is a commonly used tool in many musical signal processing tasks, including Automatic Music Transcription (AMT). However unsupervised NMF is seen to be problematic in this context, and harmonically constrained variants of NMF have been proposed. While useful, t…

Cited by 0SourceScholar