← Search

Ha-Jin Yu

13 accepted papers

2024

Diff-SV: A Unified Hierarchical Framework for Noise-Robust Speaker Verification Using Score-Based Diffusion Probabilistic Models

ICASSP 2024accepted

Background noise considerably reduces the accuracy and reliability of speaker verification (SV) systems. These challenges can be addressed using a speech enhancement system as a front-end module. Recently, diffusion probabilistic models (DPMs) have exhibited remarkable noise-compensation capabilitie…

Cited by 0SourceScholar
2024

HM-CONFORMER: A Conformer-Based Audio Deepfake Detection System with Hierarchical Pooling and Multi-Level Classification Token Aggregation Methods

ICASSP 2024accepted

Audio deepfake detection (ADD) is the task of detecting spoofing attacks generated by text-to-speech or voice conversion systems. Spoofing evidence, which helps to distinguish between spoofed and bona-fide utterances, might exist either locally or globally in the input features. To capture these, th…

Cited by 0SourceScholar
2023

SS-BSN: Attentive Blind-Spot Network for Self-Supervised Denoising with Nonlocal Self-Similarity

IJCAI 2023poster

Recently, numerous studies have been conducted on supervised learning-based image denoising methods. However, these methods rely on large-scale noisy-clean image pairs, which are difficult to obtain in practice. Denoising methods with self-supervised training that can be trained with only noisy imag…

2022

AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks

ICASSP 2022accepted

Artefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, si…

Cited by 0SourceScholar
2022

Attentive Max Feature Map and Joint Training for Acoustic Scene Classification

ICASSP 2022accepted

Various attention mechanisms are being widely applied to acoustic scene classification. However, we empirically found that the attention mechanism can excessively discard potentially valuable information, despite improving performance. We propose the attentive max feature map that combines two effec…

Cited by 0SourceScholar
2022

Graph Attentive Feature Aggregation for Text-Independent Speaker Verification

ICASSP 2022accepted

The objective of this paper is to combine multiple frame-level features into a single utterance-level representation considering pair-wise relationships. For this purpose, we propose a novel graph attentive feature aggregation module by interpreting each frame-level feature as a node of a graph. The…

Cited by 0SourceScholar
2022

RawNeXt: Speaker Verification System For Variable-Duration Utterances With Deep Layer Aggregation And Extended Dynamic Scaling Policies

ICASSP 2022accepted

Despite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker verification system called RawNeXt that can handle input raw waveform…

Cited by 0SourceScholar
2021

DCASENET: An Integrated Pretrained Deep Neural Network for Detecting and Classifying Acoustic Scenes and Events

ICASSP 2021accepted

Although acoustic scenes and events include many related tasks, their combined detection and classification have been scarcely investigated. We propose three architectures of deep neural networks that are integrated to simultaneously perform acoustic scene classification, audio tagging, and sound ev…

Cited by 0SourceScholar
2018

A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification Result

ICASSP 2018accepted

End-to-end systems using deep neural networks have been widely studied in the field of speaker verification. Raw audio signal processing has also been widely studied in the fields of automatic music tagging and speech recognition. However, as far as we know, end-to-end systems using raw audio signal…

Cited by 62SourceScholar
2017

Applying compensation techniques on i-vectors extracted from short-test utterances for speaker verification using deep neural network

ICASSP 2017accepted

We propose a method to improve speaker verification performance when a test utterance is very short. In some situations with short test utterances, performance of ivector/probabilistic linear discriminant analysis systems degrades. The proposed method transforms short-utterance feature vectors to ad…

Cited by 0SourceScholar
2016

Advanced b-vector system based deep neural network as classifier for speaker verification

ICASSP 2016accepted

Few studies on speaker verification have directly used a deep neural network (DNN) as a classifier. It is difficult to directly apply a DNN as a discriminative model to speaker-verification tasks because the training data for each speaker are very limited. Therefore, a b-vector has been proposed to…

Cited by 0SourceScholar