← Search

Ian McLoughlin

26 accepted papers

2025

PNP-RKD: A Positive-Negative Pair based Relational Knowledge Distillation Method for Cross-Domain Speaker Verification

ICASSP 2025accepted

Existing deep embedding learning based speaker verification (SV) methods suffer from performance degradation under domain shift conditions. This can be alleviated through unsupervised domain adaptation (UDA) techniques. While UDA improves global statistical consistency across domains, discriminative…

Cited by 0SourceScholar
2025

Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection

ICASSP 2025accepted

A significant challenge in sound event detection (SED) is the effective utilization of unlabeled data, given the limited availability of labeled data due to high annotation costs. Semi-supervised algorithms rely on labeled data to learn from unlabeled data, and the performance is constrained by the…

Cited by 0SourceScholar
2024

Meta Representation Learning Method for Robust Speaker Verification in Unseen Domains

ICASSP 2024accepted

This paper presents a meta representation learning method for robust speaker verification (SV) in unseen domains. It is known that the existing embedding learning based SV systems may suffer from domain mismatch issues. To address this, we propose an episodic training procedure to compensate domain…

Cited by 0SourceScholar
2023

AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer

ICASSP 2023accepted

In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have recently shown promise on DCASE2022 challenge task4 where t…

Cited by 0SourceScholar
2023

An Effective Anomalous Sound Detection Method Based on Representation Learning with Simulated Anomalies

ICASSP 2023accepted

In this paper, we propose an effective anomalous sound detection (ASD) method based on representation learning with simulated anomalies. Recently, ASD systems have used Outlier Exposure (OE) strategy to achieve promising performance in DCASE challenges. These exploit deep Convolutional Neural Networ…

Cited by 0SourceScholar
2023

Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection

ICASSP 2023accepted

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together wi…

Cited by 28SourceScholar
2023

Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to…

Cited by 0SourceScholar
2022

Domain Robust Deep Embedding Learning for Speaker Recognition

ICASSP 2022accepted

This paper presents a domain robust deep embedding learning method for speaker verification (SV) tasks. Most recent methods utilize deep neural networks (DNN) to learn compact and discriminative speaker embeddings from large-scale labeled datasets such as VoxCeleb and the NIST SRE corpus. Despite th…

Cited by 0SourceScholar
2022

Frontend Attributes Disentanglement for Speech Emotion Recognition

ICASSP 2022accepted

Speech emotion recognition (SER) with limited size dataset is a challenging task, since a spoken utterance contains various disturbing attributes besides emotion, including speaker, content, and language. However, due to a close relationship between speaker and emotion attributes, simply fine-tuning…

Cited by 0SourceScholar
2022

Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift

ICASSP 2022accepted

In this paper, a self-supervised representation learning method is proposed for anomalous sound detection (ASD). ASD has received much research attention in recent DCASE challenges. It aims to identify whether a sound emitted from a machine is anomalous or not, given only normal sound data. This is…

Cited by 0SourceScholar
2021

An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification

ICASSP 2021accepted

In this paper, we present an effective end-to-end deep embedding learning method based on Dense-Residual networks, which combine the advantages of a densely connected convolutional network (DenseNet) and a residual network (ResNet), for speaker verification (SV). Unlike a model ensemble strategy whi…

Cited by 0SourceScholar
2021

An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection

ICASSP 2021accepted

This paper presents an improved mean teacher (MT) based method for large-scale weakly labeled semi-supervised sound event detection (SED), by focusing on learning a better student model. Two main improvements are proposed based on the authors’ previous perturbation based MT method. Firstly, an event…

Cited by 26SourceScholar
2021

Multi-View Audio And Music Classification

ICASSP 2021accepted

We propose in this work a multi-view learning approach for audio and music classification. Considering four typical low-level representations (i.e. different views) commonly used for audio and music recognition tasks, the proposed multi-view network consists of four subnetworks, each handling one in…

Cited by 19SourceScholar
2021

Self-Attention Generative Adversarial Network for Speech Enhancement

ICASSP 2021accepted

Existing generative adversarial networks (GANs) for speech enhancement solely rely on the convolution operation, which may obscure temporal dependencies across the sequence input. To remedy this issue, we propose a self-attention layer adapted from non-local attention, coupled with the convolutional…

Cited by 0SourceScholar
2020

An Online Speaker-aware Speech Separation Approach Based on Time-domain Representation

ICASSP 2020accepted

Despite the significant progress of deep learning based speech separation methods, it remains challenging to extract and track the speech from target speakers, especially in a single-channel multiple speaker situation. Previously, the authors proposed a source-aware context network to exploit the te…

Cited by 0SourceScholar
2020

Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection

ICASSP 2020accepted

Weakly labeled semi-supervised learning methods have recently drawn increasing attention from the research community for sound event detection tasks. Due to the weakness of the labelling, neural networks are often designed to perform sound event detection (SED) and audio tagging (AT) at the same tim…

Cited by 0SourceScholar
2019

A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification

ICASSP 2019accepted

Recently, an attention based convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) has achieved state-of-the-art performance for audio tagging (AT) and sound event detection (SED) tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challen…

Cited by 0SourceScholar
2019

Unifying Isolated and Overlapping Audio Event Detection with Multi-label Multi-task Convolutional Recurrent Neural Networks

ICASSP 2019accepted

We propose a multi-label multi-task framework based on a convolutional recurrent neural network to unify detection of isolated and overlapping audio events. The framework leverages the power of convolutional recurrent neural network architectures; convolutional layers learn effective features over w…

Cited by 22SourceScholar
2018

Source-Aware Context Network for Single-Channel Multi-Speaker Speech Separation

ICASSP 2018accepted

Deep learning based approaches have achieved promising performance in speaker-dependent single-channel multispeaker speech separation. However, partly due to the label permutation problem, they may encounter difficulties in speaker-independent conditions. Recent methods address this problem by some…

Cited by 0SourceScholar
2016

Compact convolutional neural network transfer learning for small-scale image classification

ICASSP 2016accepted

Transfer learning methods have demonstrated state-of-the-art performance on various small-scale image classification tasks. This is generally achieved by exploiting the information from an ImageNet convolution neural network (ImageNet CNN). However, the transferred CNN model is generally with high c…

Cited by 0SourceScholar
2016

Learning compact structural representations for audio events using regressor banks

ICASSP 2016accepted

We introduce a new learned descriptor for audio signals which is efficient for event representation. The entries of the descriptor are produced by evaluating a set of regressors on the input signal. The regressors are class-specific and trained using the random regression forests framework. Given an…

Cited by 0SourceScholar
2015

Improved language identification using deep bottleneck network

ICASSP 2015accepted

Effective representation plays an important role in automatic spoken language identification (LID). Recently, several representations that employ a pre-trained deep neural network (DNN) as the front-end feature extractor, have achieved state-of-the-art performance. However the performance is still f…

Cited by 0SourceScholar
2015

Multi-task deep neural network acoustic models with model adaptation using discriminative speaker identity for whisper recognition

ICASSP 2015accepted

This paper presents a study on large vocabulary continuous whisper automatic recognition (wLVCSR). wLVCSR provides the ability to use ASR equipment in public places without concern for disturbing others or leaking private information. However the task of wLVCSR is much more challenging than normal L…

Cited by 0SourceScholar