← Search

KiHyun Nam

4 accepted papers

2026

DIFFUSION-LINK: DIFFUSION PROBABILISTIC MODEL FOR BRIDGING THE AUDIO-TEXT MODALITY GAP

ICASSP 2026poster

Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Diffusion-Link, a diffusion-based modality-bridging module that generatively maps a…

Cited by 0SourcePDFScholar
2024

Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification

ICASSP 2024accepted

In the field of speaker verification, session or channel variability poses a significant challenge. While many contemporary methods aim to disentangle session information from speaker embeddings, we introduce a novel approach using an additional embedding to represent the session information. This i…

Cited by 0SourceScholar
2024

TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning

ICASSP 2024accepted

The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures while learning effective representations has been less explored. In this work, we p…

Cited by 0SourceScholar
2024

VoxMM: Rich Transcription of Conversations in the Wild

ICASSP 2024accepted

This paper presents a multi-modal dataset that contains rich transcriptions of spoken conversations. As diverse multi-modal and multi-task models emerge, there is a growing need for multi-modal training and evaluation datasets accompanied by rich metadata. However, there is no universal dataset that…

Cited by 0SourceScholar