← Search

Jaeyoung Kim

12 accepted papers

2025

Identifying and Mitigating Mismatched Language Code in Multilingual ASR

ICASSP 2025accepted

Multilingual speech recognition systems often use an input language code in order to prompt the transcription in the target language. However, the spoken language in the input audio may not always match the language code, as often prevalent in multilingual societies. This language mismatch can signi…

Cited by 0SourceScholar
2025

Query-focused Referentiability Learning for Zero-shot Retrieval

NAACL 2025long

Dense passage retrieval enhances Information Retrieval (IR) by encoding queries and passages into representation space. However, passage representations often fail to be referenced by their gold queries under domain shifts, revealing a weakness in representation space. One desirable concept for repr…

2024

HIL: Hybrid Isotropy Learning for Zero-shot Performance in Dense retrieval

NAACL 2024long

Advancements in dense retrieval models have brought ColBERT to prominence in Information Retrieval (IR) with its advanced interaction techniques.However, ColBERT is reported to frequently underperform in zero-shot scenarios, where traditional techniques such as BM25 still exceed it.Addressing this,…

2024

Monte Carlo Self-Training for Speech Recognition

ICASSP 2024accepted

Self-training in the teacher-student framework generally suffers from the confirmation bias problem, where errors from the teacher are propagated to the student and hence get amplified with multiple iterations. In this paper, we present Monte Carlo Self-training where pseudo labels are generated by…

Cited by 0SourceScholar
2023

Cross-Training: A Semi-Supervised Training Scheme for Speech Recognition

ICASSP 2023accepted

Semi-supervised training can be performed by jointly optimizing supervised and unsupervised losses. In many settings, supervised and unsupervised losses are inconsistent, and this inconsistency creates instability in training. As a solution, we propose cross-training: instead of training one network…

Cited by 3SourceScholar
2023

Key Feature Replacement of In-Distribution Samples for Out-of-Distribution Detection

AAAI 2023technical

Out-of-distribution (OOD) detection can be used in deep learning-based applications to reject outlier samples from being unreliably classified by deep neural networks. Learning to classify between OOD and in-distribution samples is difficult because data comprising the former is extremely diverse. I…

2023

Nonparametric Decoding for Generative Retrieval

ACL 2023findings

The generative retrieval model depends solely on the information encoded in its model parameters without external memory, its information capacity is limited and fixed. To overcome the limitation, we propose Nonparametric Decoding (Np Decoding) which can be applied to existing generative retrieval m…

2023

Pseudo Outlier Exposure for Out-of-Distribution Detection using Pretrained Transformers

ACL 2023findings

For real-world language applications, detecting an out-of-distribution (OOD) sample is helpful to alert users or reject such unreliable samples. However, modern over-parameterized language models often produce overconfident predictions for both in-distribution (ID) and OOD samples. In particular, la…

Cited by 3SourcePDFScholar
2022

Contrastive Siamese Network for Semi-Supervised Speech Recognition

ICASSP 2022accepted

This paper introduces contrastive siamese (c-siam) network, an architecture for leveraging unlabeled acoustic data in speech recognition. c-siam is the first network that extracts high-level linguistic information from speech by matching outputs of two identical transformer encoders. It contains aug…

Cited by 17SourceScholar
2022

Extracting Statistical Signatures of Geometry and Structure in 2D Occupancy Grid Maps for Global Localization

RA-L 2022

Global localization (or place recognition) is a method of finding the current location of a robot on a map generated by a mapping process, and it is an open field that has not yet been completely solved in the field of mobile robotics. Most existing approaches to global localization are based on ext

Cited by 12SourceScholar
2020

T-GSA: Transformer with Gaussian-Weighted Self-Attention for Speech Enhancement

ICASSP 2020accepted

Transformer neural networks (TNN) demonstrated state-ofart performance on many natural language processing (NLP) tasks, replacing recurrent neural networks (RNNs), such as LSTMs or GRUs. However, TNNs did not perform well in speech enhancement, whose contextual nature is different than NLP tasks, li…

Cited by 0SourceScholar
2018

Bridgenets: Student-Teacher Transfer Learning Based on Recursive Neural Networks and Its Application to Distant Speech Recognition

ICASSP 2018accepted

Despite the remarkable progress achieved on automatic speech recognition, recognizing far-field speeches mixed with various noise sources is still a challenging task. In this paper, we introduce novel student-teacher transfer learning, BridgeNet which can provide a solution to improve distant speech…

Cited by 0SourceScholar