← Search

Hee-Soo Heo

12 accepted papers

2024

Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification

ICASSP 2024accepted

In the field of speaker verification, session or channel variability poses a significant challenge. While many contemporary methods aim to disentangle session information from speaker embeddings, we introduce a novel approach using an additional embedding to represent the session information. This i…

Cited by 0SourceScholar
2023

Absolute Decision Corrupts Absolutely: Conservative Online Speaker Diarisation

ICASSP 2023accepted

Our focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few misjudgements in the early phase of an input session can lead to catastrophic r…

Cited by 7SourceScholar
2023

Advancing the Dimensionality Reduction of Speaker Embeddings for Speaker Diarisation: Disentangling Noise and Informing Speech Activity

ICASSP 2023accepted

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as noise, adversely affecting performance. Our previous work has…

Cited by 0SourceScholar
2023

High-Resolution Embedding Extractor for Speaker Diarisation

ICASSP 2023accepted

Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the sliding window approach, a segment easily includes two or more speakers owing to spe…

Cited by 0SourceScholar
2023

In Search of Strong Embedding Extractors for Speaker Diarisation

ICASSP 2023accepted

Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisation, from which we tackle two key problems. First, the evaluation is not straightf…

Cited by 0SourceScholar
2022

AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks

ICASSP 2022accepted

Artefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, si…

Cited by 0SourceScholar
2022

Multi-Scale Speaker Embedding-Based Graph Attention Networks For Speaker Diarisation

ICASSP 2022accepted

The objective of this work is effective speaker diarisation using multi-scale speaker embeddings. Typically, there is a trade-off between the ability to recognise short speaker segments and the discriminative power of the embedding, according to the segment length used for embedding extraction. To t…

Cited by 0SourceScholar
2021

The ins and outs of speaker recognition: lessons from VoxSRC 2020

ICASSP 2021accepted

The VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020 offers a challenging evaluation for speaker recognition systems, which includes celebrities playing different parts in movies. The goal of this work is robust speaker recognition of utterances recorded in these challenging envir…

Cited by 0SourceScholar
2018

A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification Result

ICASSP 2018accepted

End-to-end systems using deep neural networks have been widely studied in the field of speaker verification. Raw audio signal processing has also been widely studied in the fields of automatic music tagging and speech recognition. However, as far as we know, end-to-end systems using raw audio signal…

Cited by 62SourceScholar
2017

Applying compensation techniques on i-vectors extracted from short-test utterances for speaker verification using deep neural network

ICASSP 2017accepted

We propose a method to improve speaker verification performance when a test utterance is very short. In some situations with short test utterances, performance of ivector/probabilistic linear discriminant analysis systems degrades. The proposed method transforms short-utterance feature vectors to ad…

Cited by 0SourceScholar
2016

Advanced b-vector system based deep neural network as classifier for speaker verification

ICASSP 2016accepted

Few studies on speaker verification have directly used a deep neural network (DNN) as a classifier. It is difficult to directly apply a DNN as a discriminative model to speaker-verification tasks because the training data for each speaker are very limited. Therefore, a b-vector has been proposed to…

Cited by 0SourceScholar