← Search

Shawn Hershey

6 accepted papers

2023

Dataset Balancing Can Hurt Model Performance

ICASSP 2023accepted

Machine learning from training data with a skewed distribution of examples per class can lead to models that favor performance on common classes at the expense of performance on rare ones. AudioSet has a very wide range of priors over its 527 sound event classes. Classification performance on AudioS…

Cited by 0SourceScholar
2021

Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

ICLR 2021poster

Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding. However, extracting sound sources which are apparent in natural videos remains an open problem. In this work, we present AudioScope, a novel audio-visual sound separation framework that can…

Cited by 86SourcePDFScholar
2021

The Benefit of Temporally-Strong Labels in Audio Event Classification

ICASSP 2021accepted

To reveal the importance of temporal precision in ground truth audio event labels, we collected precise (∼0.1 sec resolution) "strong" labels for a portion of the AudioSet dataset. We devised a temporally-strong evaluation set (including explicit negatives of varying difficulty) and a small strong-l…

Cited by 160SourceScholar
2020

Coincidence, Categorization, and Consolidation: Learning to Recognize Sounds with Minimal Supervision

ICASSP 2020accepted

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on multimodal unsupervised learning (as infants) and active learnin…

Cited by 0SourceScholar
2018

Unsupervised Learning of Semantic Audio Representations

ICASSP 2018accepted

Even in the absence of any explicit semantic annotation, vast collections of audio recordings provide valuable information for learning the categorical structure of sounds. We consider several class-agnostic semantic constraints that apply to unlabeled nonspeech audio: (i) noise and translations in…

Cited by 0SourceScholar
2017

CNN architectures for large-scale audio classification

ICASSP 2017accepted

Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Ne…

Cited by 3037SourceScholar