← Search

Yongil Kim

12 accepted papers

2025

Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation

NAACL 2025long

In line with the principle of honesty, there has been a growing effort to train large language models (LLMs) to generate outputs containing epistemic markers. However, evaluation in the presence of epistemic markers has been largely overlooked, raising a critical question: Could the use of epistemic…

2025

Can You Trick the Grader? Adversarial Persuasion of LLM Judges

EMNLP 2025

As large language models (LLMs) take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly high scores? This study is the first to reveal that strategically embedded persuasive language can bias LLM judges

Cited by 0SourcePDFScholar
2025

Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

EMNLP 2025

Recently, large vision–language models (LVLMs) have emerged as the preferred tools for judging text–image alignment, yet their robustness along the visual modality remains underexplored. This work is the first study to address a key research question: Can adversarial visual manipulations systematica

Cited by 0SourcePDFScholar
2025

Ko-LongRAG: A Korean Long-Context RAG Benchmark Built with a Retrieval-Free Approach

EMNLP 2025

The rapid advancement of large language models (LLMs) significantly enhances long-context Retrieval-Augmented Generation (RAG), yet existing benchmarks focus primarily on English. This leaves low-resource languages without comprehensive evaluation frameworks, limiting their progress in retrieval-bas

2025

LLMs can be easily Confused by Instructional Distractions

ACL 2025long

Despite the fact that large language models (LLMs) show exceptional skill in instruction following tasks, this strength can turn into a vulnerability when the models are required to disregard certain instructions. Instruction following tasks typically involve a clear task description and input text…

Cited by 0SourcePDFScholar
2025

Reasoning Models Better Express Their Confidence

NeurIPS 2025poster

Despite their strengths, large language models (LLMs) often fail to communicate their confidence accurately, making it difficult to assess when they might be wrong and limiting their reliability. In this work, we demonstrate that reasoning models that engage in extended chain-of-thought (CoT) reason…

Cited by 0SourcecodeScholar
2025

SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models

NAACL 2025findings

Despite the success of Large Language Models (LLMs), they still face challenges related to high inference costs and memory requirements. To address these issues, Knowledge Distillation (KD) has emerged as a popular method for model compression, with the use of student-generated outputs (SGOs) as tra…

2024

Kosmic: Korean Text Similarity Metric Reflecting Honorific Distinctions

COLING 2024main

Existing English-based text similarity measurements primarily focus on the semantic dimension, neglecting the unique linguistic attributes found in languages like Korean, where honorific expressions are explicitly integrated. To address this limitation, this study proposes Kosmic, a novel Korean tex…

Cited by 0SourcePDFScholar
2024

MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs

EMNLP 2024main

Despite advancements in on-topic dialogue systems, effectively managing topic shifts within dialogues remains a persistent challenge, largely attributed to the limited availability of training datasets. To address this issue, we propose Multi-Passage to Dialogue (MP2D), a data generation framework t…

Cited by 0SourcePDFScholar
2023

Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual Sources

EMNLP 2023long main

To address the data scarcity issue in Conversational question answering (ConvQA), a dialog inpainting method, which utilizes documents to generate ConvQA datasets, has been proposed. However, the original dialog inpainting model is trained solely on the dialog reconstruction task, resulting in the g…

Cited by 0SourceScholar
2023

Injecting Comparison Skills in Task-Oriented Dialogue Systems for Database Search Results Disambiguation

ACL 2023findings

In task-oriented dialogue (TOD) systems designed to aid users accomplish specific goals in one or more domains, the agent retrieves entities that satisfy user constraints from the database. However, when multiple database search results exist, an ambiguity occurs regarding which results to select an…

2023

PR-MCS: Perturbation Robust Metric for MultiLingual Image Captioning

EMNLP 2023long findings

Vulnerability to lexical perturbation is a critical weakness of automatic evaluation metrics for image captioning. This paper proposes Perturbation Robust Multi-Lingual CLIPScore(PR-MCS), which exhibits robustness to such perturbations, as a novel reference-free image captioning metric applicable to…

Cited by 0SourceScholar