← Search

Yin Cao

6 accepted papers

2022

A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection

ICASSP 2022accepted

Polyphonic sound event localization and detection (SELD) aims at detecting types of sound events with corresponding temporal activities and spatial locations. In this paper, a trackwise ensemble event independent network with a novel data augmentation method is proposed. The proposed model is based…

Cited by 0SourceScholar
2021

An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection

ICASSP 2021accepted

Polyphonic sound event localization and detection (SELD), which jointly performs sound event detection (SED) and direction-of-arrival (DoA) estimation, detects the type and occurrence time of sound events as well as their corresponding DoA angles simultaneously. We study the SELD task from a multi-t…

Cited by 0SourceScholar
2020

Learning With Out-of-Distribution Data for Audio Classification

ICASSP 2020accepted

In supervised machine learning, the assumption that training data is labelled correctly is not always satisfied. In this paper, we investigate an instance of labelling error for classification tasks in which the dataset is corrupted with out-of-distribution (OOD) instances: data that does not belong…

Cited by 0SourceScholar
2020

Source Separation with Weakly Labelled Data: an Approach to Computational Auditory Scene Analysis

ICASSP 2020accepted

Source separation is the task of separating an audio recording into individual sound sources. Source separation is fundamental for computational auditory scene analysis. Previous work on source separation has focused on separating particular sound classes such as speech and music. Much previous work…

Cited by 0SourceScholar
2019

Acoustic Scene Generation with Conditional Samplernn

ICASSP 2019accepted

Acoustic scene generation (ASG) is a task to generate waveforms for acoustic scenes. ASG can be used to generate audio scenes for movies and computer games. Recently, neural networks such as SampleRNN have been used for speech and music generation. However, ASG is more challenging due to its wide va…

Cited by 0SourceScholar
2019

Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and Challenge

ICASSP 2019accepted

Humans are able to identify a large number of environmental sounds and categorise them according to high-level semantic categories, e.g. urban sounds or music. They are also capable of generalising from past experience to new sounds when applying these categories. In this paper we report on the crea…

Cited by 14SourceScholar