← Search

Kangwook Jang

4 accepted papers

2025

Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis

ICASSP 2025accepted

This paper examines how linguistic similarity affects cross-lingual phonetic representation in speech processing for low-resource languages, emphasizing effective source language selection. Previous cross-lingual research has used various source languages to enhance performance for the target low-re…

Cited by 0SourceScholar
2025

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

ICML 2025poster

Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up without compromising computational efficiency. In this study, we introduce MoHAVE…

Cited by 2SourcePDFScholar
2025

Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation

ICLR 2025poster

Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely addressed audio disruptions, few studies have dealt with visual…

2024

STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models

ICASSP 2024accepted

Albeit great performance of Transformer-based speech self-supervised learning (SSL) models, their large parameter size and computational cost make them unfavorable to utilize. In this study, we propose to compress the speech SSL models by distilling speech temporal relation (STaR). Unlike previous w…

Cited by 0SourceScholar