← Search

Bugeun Kim

7 accepted papers

2026

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

ICML 2026poster

While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and reliable measurement instruments. To address this limitation, we introduce…

Cited by 0SourceScholar
2025

Leveraging Large Language Models for Active Merchant Non-player Characters

IJCAI 2025

We highlight two significant issues leading to the passivity of current merchant non-player characters (NPCs): pricing and communication. While immersive interactions with active NPCs have been a focus, price negotiations between merchant NPCs and players remain underexplored. First, passive pricing

2025

People will agree what I think: Investigating LLM’s False Consensus Effect

NAACL 2025findings

Large Language Models (LLMs) have been recently adopted in interactive systems requiring communication. As the false belief in a model can harm the usability of such systems, LLMs should not have cognitive biases that humans have. Psychologists especially focus on the False Consensus Effect (FCE), a…

2025

VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model

EMNLP 2025

We introduce VoiceBBQ, a spoken extension of the BBQ (Bias Benchmark for Question answering) - a dataset that measures social bias by presenting ambiguous or disambiguated contexts followed by questions that may elicit stereotypical responses. Due to the nature of speech modality, social bias in Spo

2022

EPT-X: An Expression-Pointer Transformer model that generates eXplanations for numbers

ACL 2022long

In this paper, we propose a neural model EPT-X (Expression-Pointer Transformer with Explanations), which utilizes natural language explanations to solve an algebraic word problem. To enhance the explainability of the encoding process of a neural model, EPT-X adopts the concepts of plausibility and f…