← Search

Taegwan Kang

6 accepted papers

2025

Can You Trick the Grader? Adversarial Persuasion of LLM Judges

EMNLP 2025

As large language models (LLMs) take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly high scores? This study is the first to reveal that strategically embedded persuasive language can bias LLM judges

Cited by 0SourcePDFScholar
2025

Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

EMNLP 2025

Recently, large vision–language models (LVLMs) have emerged as the preferred tools for judging text–image alignment, yet their robustness along the visual modality remains underexplored. This work is the first study to address a key research question: Can adversarial visual manipulations systematica

Cited by 0SourcePDFScholar
2025

LLMs can be easily Confused by Instructional Distractions

ACL 2025long

Despite the fact that large language models (LLMs) show exceptional skill in instruction following tasks, this strength can turn into a vulnerability when the models are required to disregard certain instructions. Instruction following tasks typically involve a clear task description and input text…

Cited by 0SourcePDFScholar
2025

SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models

NAACL 2025findings

Despite the success of Large Language Models (LLMs), they still face challenges related to high inference costs and memory requirements. To address these issues, Knowledge Distillation (KD) has emerged as a popular method for model compression, with the use of student-generated outputs (SGOs) as tra…

2024

Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination

ACL 2024long

Instruction-following language models often show undesirable biases. These undesirable biases may be accelerated in the real-world usage of language models, where a wide range of instructions is used through zero-shot example prompting. To solve this problem, we first define the bias neuron, which s…

Cited by 7SourcePDFScholar
2022

Deriving Explainable Discriminative Attributes Using Confusion About Counterfactual Class

ICASSP 2022accepted

Recently, Integrated Gradients-based (IG) methods have been commonly used to explain the decision process of deep neural networks (DNNs). However, they have only considered the information of the predicted class while neglecting the in-formation of the rest classes. In this paper, we propose a novel…

Cited by 0SourceScholar