← Search

Joon-Hyuk Chang

26 accepted papers

2026

OnEDIT: Online Editing with Decoupled Implicit Task for Large Language Models

AAAI 2026technical

Continual instruction tuning (CIT) has emerged as a promising strategy for adapting large language models (LLMs) to new tasks while preserving historical knowledge. Most existing CIT methods have focused on offline CIT (offCIT), which assumes clearly defined task boundaries and allows multiple passe

Cited by 0SourcePDFScholar
2026

Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization

ICLR 2026poster

Sound source localization (SSL) is a fundamental task in spatial audio understanding, yet most deep neural network-based methods are constrained by fixed array geometries and predefined directional grids, limiting generalizability and scalability. To address these issues, we propose _audio-geometry-…

Cited by 0SourcecodeScholar
2025

Diffusion-based Target Device Style Transfer for Robust Acoustic Scene Classification

ICASSP 2025accepted

Audio signal processing systems often operate differently depending on the recording devices, leading to performance discrepancies. Therefore, it is important to know about the characteristics of the recording device; however, it is difficult to know the device’s behavior in most cases. In this stud…

Cited by 0SourceScholar
2025

Improving Target Sound Extraction via Disentangled Codec Representations with Privileged Knowledge Distillation

NeurIPS 2025poster

Target sound extraction aims to isolate target sound sources from an input mixture using a target clue to identify the sounds of interest. To address the challenge posed by the wide variety of sounds, recent work has introduced privileged knowledge distillation (PKD), which utilizes privileged infor…

Cited by 0SourceScholar
2025

Quad-Net: Melspectrogram Vocoder with Convolutional Layers Restricted by the Quadrature Mirror Filter for Perfect Reconstruction

ICASSP 2025accepted

Recently, neural vocoders have applied signal processing methods to synthesize speech to reduce computational complexity. However, most methods lack the benefits of a data-driven approach and the flexibility of hyper-parameters, such as filter length, because they rely on fixed signal processing fil…

Cited by 0SourceScholar
2025

Trainable Adaptive Score Normalization for Automatic Speaker Verification

ICASSP 2025accepted

Adaptive S-norm (AS-norm) calibrates automatic speaker verification (ASV) scores by normalizing them utilize the scores of impostors which are similar to the input speaker. However, AS-norm does not involve any learning process, limiting its ability to provide appropriate regularization strength for…

Cited by 0SourceScholar
2024

Adversarial Learning on Compressed Posterior Space for Non-Iterative Score-based End-to-End Text-to-Speech

ICASSP 2024accepted

Score-based generative models have shown the real-like quality of synthesized speech in the text-to-speech (TTS) area. However, the critical artifact of score-based models is the requirement of a high computational cost due to the iterative sampling algorithm, and it also makes it difficult to fine-…

Cited by 0SourceScholar
2024

Generalized Specaugment via Multi-Rectangle Inverse Masking For Acoustic Scene Classification

ICASSP 2024accepted

In this paper, we present the multi-rectangle inverse masking (MRIM), an extension and generalization of the traditional SpecAugment technique, for acoustic scene classification. While SpecAugment, observed from its unmasked areas, primarily forms rectangles around the input corners, our novel strat…

Cited by 0SourceScholar
2024

Improving Target Sound Extraction with Timestamp Knowledge Distillation

ICASSP 2024accepted

In this paper, we propose a timestamp knowledge distillation (TKD) method that adopts privileged knowledge distillation to enhance the performance of deep neural network (DNN)-based target sound extraction (TSE). While previous studies have mainly used n-hot vectors to indicate the type of target so…

Cited by 0SourceScholar
2024

Stationary Latent Weight Inference for Unreliable Observations from Online Test-Time Adaptation

ICML 2024poster

In the rapidly evolving field of online test-time adaptation (OTTA), effectively managing distribution shifts is a pivotal concern. State-of-the-art OTTA methodologies often face limitations such as an inadequate target domain information integration, leading to significant issues like catastrophic…

Cited by 2SourcePDFScholar
2024

Text-Only Unsupervised Domain Adaptation for Neural Transducer-Based ASR Personalization Using Synthesized Data

ICASSP 2024accepted

Research on personalizing neural transducer-based automatic speech recognition (ASR) systems using the text-only data is currently flourishing. Among various approaches, utilizing synthesized speech offers an advantage of adapting the entire ASR system. In this study, we explore the problem of perso…

Cited by 0SourceScholar
2023

Adaptive Time-Scale Modification for Improving Speech Intelligibility Based On Phoneme Clustering For Streaming Services

ICASSP 2023accepted

Time-scale modification (TSM) is important in streaming services, including over-the-top (OTT) platforms, audiobooks, and online lectures. Although TSM modifies the speed of audio while maintaining other audio attributes such as the pitch and timbre of the speaker, it unnaturally distorts audio sign…

Cited by 0SourceScholar
2023

Improving Transformer-Based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads

ICASSP 2023accepted

Transformer-based end-to-end neural speaker diarization (EEND) models utilize the multi-head self-attention (SA) mechanism to enable accurate speaker label prediction in overlapped speech regions. In this study, to enhance the training effectiveness of SA-EEND models, we propose the use of auxiliary…

Cited by 0SourceScholar
2023

M-CTRL: A Continual Representation Learning Framework with Slowly Improving Past Pre-Trained Model

ICASSP 2023accepted

Representation models pre-trained on unlabeled data show competitive performance in speech recognition, even when fine-tuned on small amounts of labeled data. The continual representation learning (CTRL) framework combines pre-training and continual learning methods to obtain powerful representation…

Cited by 0SourceScholar
2023

Noise-Aware Target Extension with Self-Distillation for Robust Speech Recognition

ICASSP 2023accepted

Data augmentation using additive noise is a framework for robustly training automatic speech recognition models. To utilize noise information efficiently, previous studies used an additional branch to classify noise conditions. This added branch has a limited effect on the ASR because it performs in…

Cited by 0SourceScholar
2023

Repackagingaugment: Overcoming Prediction Error Amplification in Weight-Averaged Speech Recognition Models Subject to Self-Training

ICASSP 2023accepted

Representation-based speech recognition models have demonstrated state-of-the-art performance on downstream tasks. These models are pre-trained on large-scale unlabeled data, fine-tuned on a small amount of labeled data, and subsequently advanced via the self-training procedure by leveraging pseudo-…

Cited by 0SourceScholar
2022

Knowledge Distillation from Language Model to Acoustic Model: A Hierarchical Multi-Task Learning Approach

ICASSP 2022accepted

The remarkable performance of the pre-trained language model (LM) using self-supervised learning has led to a major paradigm shift in the study of natural language processing. In line with these changes, leveraging the performance of speech recognition systems with massive deep learning-based LMs is…

Cited by 0SourceScholar
2018

DNN-based Speech Recognition System dealing with Motor State as Auxiliary Information of DNN for Head Shaking Robot

IROS 2018poster

In this paper, a deep neural network (DNN) based integrated background noise suppression and acoustic modeling for speech recognition proposed in which on/off state of the motor for the head shaking robot is employed as the relevant auxiliary information of the DNN input. Since the motor sound being…

Cited by 4SourceScholar
2016

Dual-microphone voice activity detection based on using optimally weighted maximum a posteriori probabilities

ICASSP 2016accepted

In this paper, we propose to improve the dual-microphone voice activity detection (VAD) technique for which a discriminative weight training is applied to achieve optimally weighted spatial features. In our approach, we first derive the maximum a posteriori (MAP) probabilities from the spatial featu…

Cited by 0SourceScholar