← Search

Jee-weon Jung

21 accepted papers

2025

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

ICASSP 2025accepted

Iterative self-training, or iterative pseudo-labeling (IPL)—using an improved model from the current iteration to provide pseudo-labels for the next iteration—has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker re…

Cited by 0SourceScholar
2024

AugSumm: Towards Generalizable Speech Summarization Using Synthetic Labels from Large Language Models

ICASSP 2024accepted

Abstractive speech summarization (SSUM) aims to generate humanlike summaries from speech. Given variations in information captured and phrasing, recordings can be summarized in multiple ways. Therefore, it is more reasonable to consider a probabilistic distribution of all potential summaries rather…

Cited by 0SourceScholar
2024

Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study

ICASSP 2024accepted

Speech signals, typically sampled at rates in the tens of thousands per second, contain redundancies, evoking inefficiencies in sequence modeling. High-dimensional speech features such as spectrograms are often used as the input for the subsequent model. However, they can still be redundant. Recent…

Cited by 0SourceScholar
2024

Improving Audio Captioning Models with Fine-Grained Audio Features, Text Embedding Supervision, and LLM Mix-Up Augmentation

ICASSP 2024accepted

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a sequence-to-sequence (seq2seq) backbone powered by strong mo…

Cited by 0SourceScholar
2024

On the Evaluation of Speech Foundation Models for Spoken Language Understanding

ACL 2024findings

The Spoken Language Understanding Evaluation (SLUE) suite of benchmark tasks was recently introduced to address the need for openresources and benchmarking of complex spoken language understanding (SLU) tasks, including both classification and sequence generation tasks, on natural speech. The benchm…

Cited by 6SourcePDFScholar
2024

One Model to Rule Them All ? Towards End-to-End Joint Speaker Diarization and Speech Recognition

ICASSP 2024accepted

This paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any number of speakers, effectively solving "who spoke what, whe…

Cited by 0SourceScholar
2024

Understanding Probe Behaviors Through Variational Bounds of Mutual Information

ICASSP 2024accepted

With the success of self-supervised representations, researchers seek a better understanding of the information encapsulated within a representation. Among various interpretability methods, we focus on classification-based linear probing. We aim to foster a solid understanding and provide guidelines…

Cited by 0SourceScholar
2024

UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions

NAACL 2024long

Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model’s behavior and surpassing performance of task-specific models. Motivated by this, we ask: can we build a single model that jointly performs various spoken language underst…

2024

VoxMM: Rich Transcription of Conversations in the Wild

ICASSP 2024accepted

This paper presents a multi-modal dataset that contains rich transcriptions of spoken conversations. As diverse multi-modal and multi-task models emerge, there is a growing need for multi-modal training and evaluation datasets accompanied by rich metadata. However, there is no universal dataset that…

Cited by 0SourceScholar
2024

VoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation Tasks

ICASSP 2024accepted

We propose a decoder-only language model, VoxtLM, that can perform four tasks: speech recognition, speech synthesis, text generation, and speech continuation. VoxtLM integrates text vocabulary with discrete speech tokens from self-supervised speech features and uses special tokens to enable multitas…

Cited by 0SourceScholar
2023

Absolute Decision Corrupts Absolutely: Conservative Online Speaker Diarisation

ICASSP 2023accepted

Our focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few misjudgements in the early phase of an input session can lead to catastrophic r…

Cited by 0SourceScholar
2023

Advancing the Dimensionality Reduction of Speaker Embeddings for Speaker Diarisation: Disentangling Noise and Informing Speech Activity

ICASSP 2023accepted

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as noise, adversely affecting performance. Our previous work has…

Cited by 0SourceScholar
2023

High-Resolution Embedding Extractor for Speaker Diarisation

ICASSP 2023accepted

Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the sliding window approach, a segment easily includes two or more speakers owing to spe…

Cited by 0SourceScholar
2023

In Search of Strong Embedding Extractors for Speaker Diarisation

ICASSP 2023accepted

Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisation, from which we tackle two key problems. First, the evaluation is not straightf…

Cited by 0SourceScholar
2022

AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks

ICASSP 2022accepted

Artefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, si…

Cited by 0SourceScholar
2022

Attentive Max Feature Map and Joint Training for Acoustic Scene Classification

ICASSP 2022accepted

Various attention mechanisms are being widely applied to acoustic scene classification. However, we empirically found that the attention mechanism can excessively discard potentially valuable information, despite improving performance. We propose the attentive max feature map that combines two effec…

Cited by 0SourceScholar
2022

Multi-Scale Speaker Embedding-Based Graph Attention Networks For Speaker Diarisation

ICASSP 2022accepted

The objective of this work is effective speaker diarisation using multi-scale speaker embeddings. Typically, there is a trade-off between the ability to recognise short speaker segments and the discriminative power of the embedding, according to the segment length used for embedding extraction. To t…

Cited by 0SourceScholar
2021

DCASENET: An Integrated Pretrained Deep Neural Network for Detecting and Classifying Acoustic Scenes and Events

ICASSP 2021accepted

Although acoustic scenes and events include many related tasks, their combined detection and classification have been scarcely investigated. We propose three architectures of deep neural networks that are integrated to simultaneously perform acoustic scene classification, audio tagging, and sound ev…

Cited by 0SourceScholar
2018

A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification Result

ICASSP 2018accepted

End-to-end systems using deep neural networks have been widely studied in the field of speaker verification. Raw audio signal processing has also been widely studied in the fields of automatic music tagging and speech recognition. However, as far as we know, end-to-end systems using raw audio signal…

Cited by 0SourceScholar