← Search

You Jin Kim

6 accepted papers

2024

Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification

ICASSP 2024accepted

In the field of speaker verification, session or channel variability poses a significant challenge. While many contemporary methods aim to disentangle session information from speaker embeddings, we introduce a novel approach using an additional embedding to represent the session information. This i…

Cited by 0SourceScholar
2024

TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning

ICASSP 2024accepted

The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures while learning effective representations has been less explored. In this work, we p…

Cited by 0SourceScholar
2023

Absolute Decision Corrupts Absolutely: Conservative Online Speaker Diarisation

ICASSP 2023accepted

Our focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few misjudgements in the early phase of an input session can lead to catastrophic r…

Cited by 7SourceScholar
2023

Advancing the Dimensionality Reduction of Speaker Embeddings for Speaker Diarisation: Disentangling Noise and Informing Speech Activity

ICASSP 2023accepted

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as noise, adversely affecting performance. Our previous work has…

Cited by 0SourceScholar
2023

High-Resolution Embedding Extractor for Speaker Diarisation

ICASSP 2023accepted

Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the sliding window approach, a segment easily includes two or more speakers owing to spe…

Cited by 0SourceScholar
2022

Multi-Scale Speaker Embedding-Based Graph Attention Networks For Speaker Diarisation

ICASSP 2022accepted

The objective of this work is effective speaker diarisation using multi-scale speaker embeddings. Typically, there is a trade-off between the ability to recognise short speaker segments and the discriminative power of the embedding, according to the segment length used for embedding extraction. To t…

Cited by 0SourceScholar