← Search

Yunsu Kim

17 accepted papers

2025

DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction

ACL 2025finding

We present DeRAGEC, a method for improving Named Entity (NE) correction in Automatic Speech Recognition (ASR) systems. By extending the Retrieval-Augmented Generative Error Correction (RAGEC) framework, DeRAGEC employs synthetic denoising rationales to filter out noisy NE candidates before correctio…

2025

DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition

NAACL 2025long

Dysarthric speech recognition often suffers from performance degradation due to the intrinsic diversity of dysarthric severity and extrinsic disparity from normal speech. To bridge these gaps, we propose a Dynamic Phoneme-level Contrastive Learning (DyPCL) method, which leads to obtaining invariant…

2025

MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queries

EMNLP 2025

Despite bilingual speakers frequently using mixed-language queries in web searches, Information Retrieval (IR) research on them remains scarce. To address this, we introduce ***MiLQ***, ***Mi***xed-***L***anguage ***Q***uery test set, the first public benchmark of mixed-language queries, qualified a

2025

Revisiting Early Detection of Sexual Predators via Turn-level Optimization

NAACL 2025long

Online grooming is a severe social threat where sexual predators gradually entrap child victims with subtle and gradual manipulation. Therefore, timely intervention for online grooming is critical for proactive protection. However, previous methods fail to determine the optimal intervention points (…

2025

sDPO: Don’t Use Your Data All at Once

COLING 2025industry

As large language models (LLMs) continue to advance, aligning them with human preferences has become a critical objective. In this paper, we introduce stepwise DPO (sDPO), an innovative extension of the recently popularized Direct Preference Optimization (DPO) technique for alignment tuning. sDPO sy…

Cited by 27SourcePDFScholar
2024

Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing

EMNLP 2024main

Recent efforts have aimed to utilize multilingual pretrained language models (mPLMs) to extend semantic parsing (SP) across multiple languages without requiring extensive annotations. However, achieving zero-shot cross-lingual transfer for SP remains challenging, leading to a performance gap between…

2024

Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target Languages

EMNLP 2024main

Automatic question generation (QG) serves a wide range of purposes, such as augmenting question-answering (QA) corpora, enhancing chatbot systems, and developing educational materials. Despite its importance, most existing datasets predominantly focus on English, resulting in a considerable gap in d…

2024

Denoising Table-Text Retrieval for Open-Domain Question Answering

COLING 2024main

In table-text open-domain question answering, a retriever system retrieves relevant evidence from tables and text to answer questions. Previous studies in table-text open-domain question answering have two common challenges: firstly, their retrievers can be affected by false-positive labels in train…

2024

Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations

AAAI 2024technical

Zero-shot multi-speaker TTS aims to synthesize speech with the voice of a chosen target speaker without any fine-tuning. Prevailing methods, however, encounter limitations at adapting to new speakers of out-of-domain settings, primarily due to inadequate speaker disentanglement and content leakage.…

2024

Evalverse: Unified and Accessible Library for Large Language Model Evaluation

EMNLP 2024system demonstrations

This paper introduces Evalverse, a novel library that streamlines the evaluation of Large Language Models (LLMs) by unifying disparate evaluation tools into a single, user-friendly framework. Evalverse enables individuals with limited knowledge of artificial intelligence to easily request LLM evalua…

2024

Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling

COLING 2024main

In response to the increasing use of interactive artificial intelligence, the demand for the capacity to handle complex questions has increased. Multi-hop question generation aims to generate complex questions that requires multi-step reasoning over several documents. Previous studies have predomina…

2024

Leveraging the Interplay between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation

COLING 2024main

Contemporary neural speech synthesis models have indeed demonstrated remarkable proficiency in synthetic speech generation as they have attained a level of quality comparable to that of human-produced speech. Nevertheless, it is important to note that these achievements have predominantly been verif…

Cited by 0SourcePDFScholar
2024

Multi-Dimensional Optimization for Text Summarization via Reinforcement Learning

ACL 2024long

The evaluation of summary quality encompasses diverse dimensions such as consistency, coherence, relevance, and fluency. However, existing summarization methods often target a specific dimension, facing challenges in generating well-balanced summaries across multiple dimensions. In this paper, we pr…

Cited by 7SourcePDFScholar
2024

SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling

NAACL 2024industry

We introduce SOLAR 10.7B, a large language model (LLM) with 10.7 billion parameters, demonstrating superior performance in various natural language processing (NLP) tasks. Inspired by recent efforts to efficiently up-scale LLMs, we present a method for scaling LLMs called depth up-scaling (DUS), whi…

2023

Bring More Attention to Syntactic Symmetry for Automatic Postediting of High-Quality Machine Translations

ACL 2023short

Automatic postediting (APE) is an automated process to refine a given machine translation (MT). Recent findings present that existing APE systems are not good at handling high-quality MTs even for a language pair with abundant data resources, English–German: the better the given MT is, the harder it…

Cited by 3SourcePDFScholar
2023

Prompt- and Trait Relation-aware Cross-prompt Essay Trait Scoring

ACL 2023findings

Automated essay scoring (AES) aims to score essays written for a given prompt, which defines the writing topic. Most existing AES systems assume to grade essays of the same prompt as used in training and assign only a holistic score. However, such settings conflict with real-education situations; pr…