← Search

Kwanghee Choi

13 accepted papers

2025

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

ICLR 2025poster

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication…

2025

Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment

NAACL 2025long

Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches…

2024

Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study

ICASSP 2024accepted

Speech signals, typically sampled at rates in the tens of thousands per second, contain redundancies, evoking inefficiencies in sequence modeling. High-dimensional speech features such as spectrograms are often used as the input for the subsequent model. However, they can still be redundant. Recent…

Cited by 0SourceScholar
2024

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

ICASSP 2024accepted

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, developed from pre-existing videos using various prediction models, and have only a small number of multi-view videos. To mitigate t…

Cited by 0SourceScholar
2024

Understanding Probe Behaviors Through Variational Bounds of Mutual Information

ICASSP 2024accepted

With the success of self-supervised representations, researchers seek a better understanding of the information encapsulated within a representation. Among various interpretability methods, we focus on classification-based linear probing. We aim to foster a solid understanding and provide guidelines…

Cited by 0SourceScholar
2024

Wav2Gloss: Generating Interlinear Glossed Text from Speech

ACL 2024long

Thousands of the world’s languages are in danger of extinction—a tremendous threat to cultural identities and human language diversity. Interlinear Glossed Text (IGT) is a form of linguistic annotation that can support documentation and resource creation for these languages’ communities. IGT typical…

2023

Automatic Severity Classification of Dysarthric Speech by Using Self-Supervised Model with Multi-Task Learning

ICASSP 2023accepted

Automatic assessment of dysarthric speech is essential for sustained treatments and rehabilitation. However, obtaining atypical speech is challenging, often leading to data scarcity issues. To tackle the problem, we propose a novel automatic severity assessment method for dysarthric speech, using th…

Cited by 0SourceScholar
2022

Combating the instability of mutual information-based losses via regularization

UAI 2022poster

Notable progress has been made in numerous fields of machine learning based on neural network-driven mutual information (MI) bounds. However, utilizing the conventional MI-based losses is often challenging due to their practical and mathematical limitations. In this work, we first identify the sympt…

2022

Learning with Noisy Labels by Efficient Transition Matrix Estimation to Combat Label Miscorrection

ECCV 2022poster

"Recent studies on learning with noisy labels have shown remarkable performance by exploiting a small clean dataset. In particular, model agnostic meta-learning-based label correction methods further improve performance by correcting noisy labels on the fly. However, there is no safeguard on the lab…

2022

Temporal Knowledge Distillation for on-device Audio Classification

ICASSP 2022accepted

Improving the performance of on-device audio classification models remains a challenge given the computational limits of the mobile environment. Many studies leverage knowledge distillation to boost predictive performance by transferring the knowledge from large models to on-device models. However,…

Cited by 0SourceScholar
2021

Disentangling Label Distribution for Long-Tailed Visual Recognition

CVPR 2021poster

The current evaluation protocol of long-tailed visual recognition trains the classification model on the long-tailed source label distribution and evaluates its performance on the uniform target label distribution. Such protocol has questionable practicality since the target may also be long-tailed.…

Cited by 312PDFcodeScholar