← Search

Daniel P. W. Ellis

14 accepted papers

2023

Dataset Balancing Can Hurt Model Performance

ICASSP 2023accepted

Machine learning from training data with a skewed distribution of examples per class can lead to models that favor performance on common classes at the expense of performance on rare ones. AudioSet has a very wide range of priors over its 527 sound event classes. Classification performance on AudioS…

Cited by 0SourceScholar
2021

The Benefit of Temporally-Strong Labels in Audio Event Classification

ICASSP 2021accepted

To reveal the importance of temporal precision in ground truth audio event labels, we collected precise (∼0.1 sec resolution) "strong" labels for a portion of the AudioSet dataset. We devised a temporally-strong evaluation set (including explicit negatives of varying difficulty) and a small strong-l…

Cited by 160SourceScholar
2021

What's all the Fuss about Free Universal Sound Separation Data?

ICASSP 2021accepted

We introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio data drawn from 357 classes, which are used to create mixtur…

Cited by 0SourceScholar
2020

Coincidence, Categorization, and Consolidation: Learning to Recognize Sounds with Minimal Supervision

ICASSP 2020accepted

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on multimodal unsupervised learning (as infants) and active learnin…

Cited by 0SourceScholar
2020

Improving Universal Sound Separation Using Sound Classification

ICASSP 2020accepted

Deep learning approaches have recently achieved impressive performance on both audio source separation and sound classification. Most audio source separation approaches focus only on separating sources belonging to a restricted domain of source classes, such as speech and music. However, recent work…

Cited by 0SourceScholar
2020

Large-Scale Weakly-Supervised Content Embeddings for Music Recommendation and Tagging

ICASSP 2020accepted

We explore content-based representation learning strategies tailored for large-scale, uncurated music collections that afford only weak supervision through unstructured natural language metadata and co-listen statistics. At the core is a hybrid training scheme that uses classification and metric lea…

Cited by 0SourceScholar
2019

Learning Sound Event Classifiers from Web Audio with Noisy Labels

ICASSP 2019accepted

As sound event classification moves towards larger datasets, issues of label noise become inevitable. Web sites can supply large volumes of user-contributed audio and metadata, but inferring labels from this metadata introduces errors due to unreliable inputs, and limitations in the mapping. There i…

Cited by 0SourceScholar
2018

Unsupervised Learning of Semantic Audio Representations

ICASSP 2018accepted

Even in the absence of any explicit semantic annotation, vast collections of audio recordings provide valuable information for learning the categorical structure of sounds. We consider several class-agnostic semantic constraints that apply to unlabeled nonspeech audio: (i) noise and translations in…

Cited by 0SourceScholar
2017

Audio Set: An ontology and human-labeled dataset for audio events

ICASSP 2017accepted

Audio event recognition, the human-like ability to identify and relate sounds from audio, is a nascent problem in machine perception. Comparable problems such as object detection in images have reaped enormous benefits from comprehensive datasets - principally ImageNet. This paper describes the crea…

Cited by 0SourceScholar
2017

CNN architectures for large-scale audio classification

ICASSP 2017accepted

Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Ne…

Cited by 3037SourceScholar
2017

Large-scale audio event discovery in one million YouTube videos

ICASSP 2017accepted

Internet videos provide a virtually boundless source of audio with a conspicuous lack of localized annotations, presenting an ideal setting for unsupervised methods. With this motivation, we perform an unprecedented exploration into the large-scale discovery of recurring audio events in a diverse co…

Cited by 0SourceScholar
2015

Micbots: Collecting large realistic datasets for speech and audio research using mobile robots

ICASSP 2015accepted

Speech and audio signal processing research is a tale of data collection efforts and evaluation campaigns. Large benchmark datasets for automatic speech recognition (ASR) have been instrumental in the advancement of speech recognition technologies. However, when it comes to robust ASR, source separa…

Cited by 0SourceScholar